Conversation
🔗 Helpful Links🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/23121
Note: Links to docs will display an error until the docs builds have been completed. ❌ 1 New Failure, 1 Pending, 1 Unclassified FailureAs of commit 813c476 with merge base 3794e44 ( NEW FAILURE - The following job has failed:
UNCLASSIFIED FAILURE - DrCI could not classify the following job because the workflow did not run on the merge base. The failure may be pre-existing on trunk or introduced by this PR:
This comment was automatically generated by Dr. CI and updates every 15 minutes. |
This PR needs a
|
162d4c7 to
b10b9e3
Compare
Stacked on #23121, which ships linkable libraries in the Windows wheel. A Windows user can run a CUDA program but cannot lower one: lowering compiles the model with a toolchain only the Linux side has, and the CUDA Windows CI job already exports on Linux and runs on Windows for exactly this reason. What was missing is a wheel that gives that Windows user the runtime. This adds a Windows CUDA wheel carrying the CUDA delegate as C++ libraries, the same components the Linux CUDA wheel ships (#21645, #21668), and nothing CUDA in the Python module, which on Windows has no CUDA program to lower. lib/executorch_backend_cuda.dll + .lib the delegate executorch::backend_cuda lib/executorch_extension_cuda.dll + .lib the stream helper executorch::extension_cuda backends/cuda/aoti_cuda_shims.dll the AOTI shim layer (runtime DLL of the delegate) data/lib/aoti_cuda_shims.lib already shipped as the lowering stub; also the shim layer's import library here extension/pybindings/_C.pyd unchanged: no CUDA dependency, no CudaBackend A C++ application links it exactly as on Linux: find_package(executorch REQUIRED COMPONENTS kernels_optimized backend_cuda) target_link_libraries(app PRIVATE executorch::runtime executorch::kernels_optimized executorch::backend_cuda) # plus the $<TARGET_RUNTIME_DLLS:app> copy from the CPU PR, which now also brings the # stream helper and the shim layer The CUDA runtime is linked into the DLLs (CMake's default on Windows), so the wheel needs the display driver and not the CUDA toolkit. nvidia publishes no CUDA runtime package for Windows, so the wheel declares none; the Linux-marked requirements are unchanged. ## How - Generator. The CUDA toolkit's Visual Studio integration fails compiler identification under the ClangCL toolset (MSB4023 in CUDA 13.0.targets), and the MSVC toolset cannot compile the Python extension's sources. A Windows CUDA build therefore uses Ninja Multi-Config with clang-cl, the compiler the CPU wheel uses, and nvcc with cl.exe as its host compiler, the combination the CUDA Windows CI job builds with. The multi-config generator keeps the Release/ layout packaging reads. The toolkit is the one install_utils already reports, so what compiles is the train packaging declares. - flatc / flatcc byproducts are declared with the executable suffix. The Visual Studio generator ignores them; Ninja on Windows had no rule producing flatc.exe. - The runtime DLL makes its platform layer strong (ET_PAL_STRONG_SYMBOLS). A DLL resolves its own references at link time so a weak PAL cannot be overridden through it anyway, and the generated export list skips weak symbols, which left the shims importing et_pal_init from nowhere. - The delegate and shims build as C++20 in the Windows shared layout, which the vendored c10 headers need on the MSVC branch; a static build gets that from executorch_core as before. The shims resolve the runtime from executorch.dll rather than a static core. - The shims keep WINDOWS_EXPORT_ALL_SYMBOLS OFF in a static build, as upstream sets it, and turn it ON in the shared layout: the delegate DLL imports the stream guard functions (get/set/peek/clearCurrentCUDAStream) from the shim DLL, and those carry no AOTI_SHIM_EXPORT, so the annotations alone would not export them. - #23121 keeps a Windows CUDA build static; this drops that guard, so the pybind preset now builds the shared layout with CUDA too, and adds the CUDA import libraries to the Windows packaging list. - extension/cuda links CUDA::cudart before backends/cuda finds the toolkit, so the root finds it first; a CPU torch does not define that target. - The package config adds the stream helper and the shim layer to the delegate's runtime DLL set, and reports EXECUTORCH_RUNTIME_DLLS_EXTRA for the pre-3.28 route, since the shim layer ships outside lib/. - CI: build-wheels-cuda-windows.yml mirrors the Linux CUDA workflow on the Windows wheel builder and reuses filter_cuda_matrix.py. pre_build_script.sh drops the blanket Windows CUDA=OFF (a CPU row is still forced off by the generic row classification) and resolves TORCH_CUDA_ARCH_LIST from cuda_arch_list.sh, since the Windows builder has no env-var script slot. The Arm Cortex-M Python module is off in that build: Ninja does not generate its two identically named sources as distinct rules. ## Tests test_cuda_windows.py, the Windows counterpart of test_cuda_linux.py: - the CUDA DLLs and import libraries ship; - _C depends on no CUDA DLL and registers no CUDA backend; - no CUDA DLL imports cudart64_*.dll (only the driver), and no Windows CUDA requirement is declared; - the shim layer's device code covers the row's TORCH_CUDA_ARCH_LIST; - a C++ application linking executorch::backend_cuda gets the three DLLs copied beside it and sees CudaBackend registered; - a program lowered on Linux for Windows (`test_cuda_windows.py --export DIR`, which the CUDA Windows CI job's export step can produce) runs through the C++ SDK and matches eager; without artifacts it prints SKIP, as the Linux aarch64 rows do for execution; - then everything a CPU Windows row checks: test_shared_libraries.py (now with the three CUDA ownership rows) and test_cpp_sdk.py. ## Test plan On Windows 11 x64 with an RTX 5080 (sm_120), VS 2022 BuildTools, CUDA 13.0, CPython 3.12: - `CU_VERSION=cu130 TORCH_CUDA_ARCH_LIST=12.0 python setup.py bdist_wheel` -> executorch-1.6.0+cu130-cp312-cp312-win_amd64.whl, installed clean. - In WSL (Ubuntu, torch 2.14.0+cu130, mingw-w64 and the Windows CUDA runtime from install_cuda_windows_cross_compile.sh), from this branch: `python .ci/scripts/wheel/test_cuda_windows.py --export C:\...\cuda-windows-artifacts` -> model.pte (4.9 MB, carrying the Windows kernel DLL) + aoti_cuda_blob.ptd. - On Windows: `EXECUTORCH_CUDA_WINDOWS_ARTIFACTS=... python test_cuda_windows.py`: every check passes, including the Linux-exported program running on the GPU through executorch::backend_cuda and matching eager to 2.4e-7. - CPU regression: the CPU wheel rebuilt from this branch (ClangCL, CUDA off) passes the full test_windows.py with the same results as #23121. - The Windows CUDA CI job (cuda-windows.yml, static llm-release-cuda build) is unaffected: every Windows change in backends/cuda is behind EXECUTORCH_BUILD_SHARED. - Linux: the CMake changes are behind WIN32 / MSVC or evaluate identically there (an empty CMAKE_EXECUTABLE_SUFFIX); setup.py's new path runs only on Windows; relying on the Linux wheel rows for confirmation.
b10b9e3 to
886edad
Compare
Stacked on #23121, which ships linkable libraries in the Windows wheel. A Windows user can run a CUDA program but cannot lower one: lowering compiles the model with a toolchain only the Linux side has, and the CUDA Windows CI job already exports on Linux and runs on Windows for exactly this reason. What was missing is a wheel that gives that Windows user the runtime. This adds a Windows CUDA wheel carrying the CUDA delegate as C++ libraries, the same components the Linux CUDA wheel ships (#21645, #21668), and nothing CUDA in the Python module, which on Windows has no CUDA program to lower. lib/executorch_backend_cuda.dll + .lib the delegate executorch::backend_cuda lib/executorch_extension_cuda.dll + .lib the stream helper executorch::extension_cuda backends/cuda/aoti_cuda_shims.dll the AOTI shim layer (runtime DLL of the delegate) data/lib/aoti_cuda_shims.lib already shipped as the lowering stub; also the shim layer's import library here extension/pybindings/_C.pyd unchanged: no CUDA dependency, no CudaBackend A C++ application links it exactly as on Linux: find_package(executorch REQUIRED COMPONENTS kernels_optimized backend_cuda) target_link_libraries(app PRIVATE executorch::runtime executorch::kernels_optimized executorch::backend_cuda) # plus the $<TARGET_RUNTIME_DLLS:app> copy from the CPU PR, which now also brings the # stream helper and the shim layer The CUDA runtime is linked into the DLLs (CMake's default on Windows), so the wheel needs the display driver and not the CUDA toolkit. nvidia publishes no CUDA runtime package for Windows, so the wheel declares none; the Linux-marked requirements are unchanged. ## How - Generator. The CUDA toolkit's Visual Studio integration fails compiler identification under the ClangCL toolset (MSB4023 in CUDA 13.0.targets), and the MSVC toolset cannot compile the Python extension's sources. A Windows CUDA build therefore uses Ninja Multi-Config with clang-cl, the compiler the CPU wheel uses, and nvcc with cl.exe as its host compiler, the combination the CUDA Windows CI job builds with. The multi-config generator keeps the Release/ layout packaging reads. The toolkit is the one install_utils already reports, so what compiles is the train packaging declares. - flatc / flatcc byproducts are declared with the executable suffix. The Visual Studio generator ignores them; Ninja on Windows had no rule producing flatc.exe. - The runtime DLL makes its platform layer strong (ET_PAL_STRONG_SYMBOLS). A DLL resolves its own references at link time so a weak PAL cannot be overridden through it anyway, and the generated export list skips weak symbols, which left the shims importing et_pal_init from nowhere. - The shims resolve the runtime from executorch.dll rather than a static core. (C++20 for the delegate and shims on Windows now comes from #23135.) - The shims keep WINDOWS_EXPORT_ALL_SYMBOLS OFF in a static build, as upstream sets it, and turn it ON in the shared layout: the delegate DLL imports the stream guard functions (get/set/peek/clearCurrentCUDAStream) from the shim DLL, and those carry no AOTI_SHIM_EXPORT, so the annotations alone would not export them. - #23121 keeps a Windows CUDA build static; this drops that guard, so the pybind preset now builds the shared layout with CUDA too, and adds the CUDA import libraries to the Windows packaging list. - extension/cuda links CUDA::cudart before backends/cuda finds the toolkit, so the root finds it first; a CPU torch does not define that target. - The package config adds the stream helper and the shim layer to the delegate's runtime DLL set, and reports EXECUTORCH_RUNTIME_DLLS_EXTRA for the pre-3.28 route, since the shim layer ships outside lib/. - CI: build-wheels-cuda-windows.yml mirrors the Linux CUDA workflow on the Windows wheel builder and reuses filter_cuda_matrix.py. pre_build_script.sh drops the blanket Windows CUDA=OFF (a CPU row is still forced off by the generic row classification) and resolves TORCH_CUDA_ARCH_LIST from cuda_arch_list.sh, since the Windows builder has no env-var script slot. The Arm Cortex-M Python module is off in that build: Ninja does not generate its two identically named sources as distinct rules. ## Tests test_cuda_windows.py, the Windows counterpart of test_cuda_linux.py: - the CUDA DLLs and import libraries ship; - _C depends on no CUDA DLL and registers no CUDA backend; - no CUDA DLL imports cudart64_*.dll (only the driver), and no Windows CUDA requirement is declared; - the shim layer's device code covers the row's TORCH_CUDA_ARCH_LIST; - a C++ application linking executorch::backend_cuda gets the three DLLs copied beside it and sees CudaBackend registered; - a program lowered on Linux for Windows (`test_cuda_windows.py --export DIR`, which the CUDA Windows CI job's export step can produce) runs through the C++ SDK and matches eager; without artifacts it prints SKIP, as the Linux aarch64 rows do for execution; - then everything a CPU Windows row checks: test_shared_libraries.py (now with the three CUDA ownership rows) and test_cpp_sdk.py. ## Test plan On Windows 11 x64 with an RTX 5080 (sm_120), VS 2022 BuildTools, CUDA 13.0, CPython 3.12: - `CU_VERSION=cu130 TORCH_CUDA_ARCH_LIST=12.0 python setup.py bdist_wheel` -> executorch-1.6.0+cu130-cp312-cp312-win_amd64.whl, installed clean. - In WSL (Ubuntu, torch 2.14.0+cu130, mingw-w64 and the Windows CUDA runtime from install_cuda_windows_cross_compile.sh), from this branch: `python .ci/scripts/wheel/test_cuda_windows.py --export C:\...\cuda-windows-artifacts` -> model.pte (4.9 MB, carrying the Windows kernel DLL) + aoti_cuda_blob.ptd. - On Windows: `EXECUTORCH_CUDA_WINDOWS_ARTIFACTS=... python test_cuda_windows.py`: every check passes, including the Linux-exported program running on the GPU through executorch::backend_cuda and matching eager to 2.4e-7. - CPU regression: the CPU wheel rebuilt from this branch (ClangCL, CUDA off) passes the full test_windows.py with the same results as #23121. - The Windows CUDA CI job (cuda-windows.yml, static llm-release-cuda build) is unaffected: every Windows change in backends/cuda is behind EXECUTORCH_BUILD_SHARED. - Linux: the CMake changes are behind WIN32 / MSVC or evaluate identically there (an empty CMAKE_EXECUTABLE_SUFFIX); setup.py's new path runs only on Windows; relying on the Linux wheel rows for confirmation.
shoumikhin
left a comment
There was a problem hiding this comment.
Some notes inline.
These are about lines the change does not touch, so they could not be attached to one.
In .ci/scripts/wheel/test_cpp_sdk.py, around line 627:
The description says every program and probe the suites run has a timeout. In the C++ suite only the runs that go through _run_consumer pass one. The runtime-only consumer run, the run without the delegate and the export step call subprocess.run with no timeout, and so do the custom op probe and the parity script in the other suite. If one of them hangs, the job waits for the CI limit instead of failing with a clear timeout. Could those calls use the same timeout, or could the sentence say which runs are bounded?
886edad to
f5e316b
Compare
24b4561 to
bd4757a
Compare
Stacked on #23121, which ships linkable libraries in the Windows wheel. A Windows user can run a CUDA program but cannot lower one: lowering compiles the model with a toolchain only the Linux side has, and the CUDA Windows CI job already exports on Linux and runs on Windows for exactly this reason. What was missing is a wheel that gives that Windows user the runtime. This adds a Windows CUDA wheel carrying the CUDA delegate as C++ libraries, the same components the Linux CUDA wheel ships (#21645, #21668), and nothing CUDA in the Python module, which on Windows has no CUDA program to lower. The export (AOT) side is unchanged and identical to the CPU wheel's; only the runtime links CUDA. lib/executorch_backend_cuda.dll + .lib the delegate executorch::backend_cuda lib/executorch_extension_cuda.dll + .lib the stream helper executorch::extension_cuda backends/cuda/aoti_cuda_shims.dll the AOTI shim layer (runtime DLL of the delegate) data/lib/aoti_cuda_shims.lib already shipped as the lowering stub; also the shim layer's import library here extension/pybindings/_C.pyd unchanged: no CUDA dependency, no CudaBackend A C++ application links it exactly as on Linux: find_package(executorch REQUIRED COMPONENTS kernels_optimized backend_cuda) target_link_libraries(app PRIVATE executorch::runtime executorch::kernels_optimized executorch::backend_cuda) # plus the $<TARGET_RUNTIME_DLLS:app> copy from the CPU PR, which now also brings the # stream helper and the shim layer The wheel's DLLs import no CUDA runtime DLL: CMake links cudart.lib into them on Windows, and that reaches the driver (nvcuda.dll). The compiled model is different: the library AOTInductor builds into model.pte imports cudart64_13.dll, which comes from the CUDA Toolkit's bin directory on PATH, as the guide says. PyPI's nvidia-cuda-runtime has win_amd64 builds that carry that DLL, but a C++ program does not search site-packages, so declaring it would not make it loadable; the wheel declares no NVIDIA package for Windows. ## How - Generator. The CUDA toolkit's Visual Studio integration fails compiler identification under the ClangCL toolset on some toolkit and Visual Studio pairs (MSB4023 in its targets file), and the MSVC toolset cannot compile the Python extension's sources. A Windows CUDA build therefore uses Ninja Multi-Config with clang-cl, the compiler the CPU wheel uses, and nvcc with cl.exe as its host compiler, when ninja is on PATH. Without it, as in a --no-build-isolation source install, the build keeps the Visual Studio generator it always used. CMAKE_CUDA_COMPILER is passed only when CUDACXX is unset, so options a user put in CUDACXX (such as -allow-unsupported-compiler) still apply. The toolkit is the one install_utils already reports, so what compiles is the train packaging declares. - CUDA 13.2 and newer. Their CCCL stops with an #error under cl.exe's traditional preprocessor, so the CUDA sources pass -Xcompiler=/Zc:preprocessor on Windows (13.0 accepts it too). - flatc / flatcc byproducts and imported locations share one host executable suffix chosen by CMAKE_HOST_WIN32. CMAKE_EXECUTABLE_SUFFIX is the target's: a Zephyr toolchain sets .elf, and naming the byproduct after it left Ninja with no rule for the host tool. - The shims resolve the runtime from executorch.dll rather than a static core. (The PAL's strong symbols are in #23121; C++20 for the delegate and shims comes from #23135.) - The shims keep WINDOWS_EXPORT_ALL_SYMBOLS OFF in a static build, as upstream sets it, and turn it ON in the shared layout: the delegate DLL imports the stream guard functions (get/set/peek/clearCurrentCUDAStream) from the shim DLL, and those carry no AOTI_SHIM_EXPORT, so the annotations alone would not export them. - #23121 keeps a Windows CUDA build static; this drops that guard, so the pybind preset now builds the shared layout with CUDA too, and adds the CUDA import libraries to the Windows packaging list. - extension/cuda links CUDA::cudart before backends/cuda finds the toolkit, so the root finds it first; a CPU torch does not define that target. - The package config adds the stream helper and the shim layer to the delegate's runtime DLL set, and reports EXECUTORCH_RUNTIME_DLLS_EXTRA for the pre-3.28 route, since the shim layer ships outside lib/. ## CI - build-wheels-cuda-windows.yml mirrors the Linux CUDA workflow on the Windows wheel builder and reuses filter_cuda_matrix.py. pre_build_script.sh drops the blanket Windows CUDA=OFF (a CPU row is still forced off by the generic row classification), resolves TORCH_CUDA_ARCH_LIST from cuda_arch_list.sh, and writes it to BUILD_ENV_FILE as well as GITHUB_ENV: the reusable workflow sources that file after the hook, and its own list had replaced the row's (sm_75 in, sm_89 out). It also installs ninja for the generator. - End-to-end, per published CUDA train, the route a Windows user has, modeled on cuda-windows.yml. The trains are read from filter_cuda_matrix.py's SUPPORTED_CUDA_VERSIONS (today cu130, cu132, cu134), the same list the release rows are built from, so adding or dropping a train there changes what is tested. (install_utils.py's list also has 12.6 for source builds, but 12.6 is no longer a published wheel train.) A Linux GPU job exports a Windows-target program with that train's torch (`test_cuda_windows.py --export`), and a Windows GPU job builds the wheel with that train's nvcc, installs it and runs test_cuda_windows.py with the program (test_cuda_windows_e2e.ps1). Trains the images do not carry are assembled from NVIDIA's checksummed redistributable archives by install_cuda_redist.py, which needs no installer or admin rights and fails on a train it has no release for. The wheel build's own smoke test has no GPU program and says so. ## Tests test_cuda_windows.py, the Windows counterpart of test_cuda_linux.py: - the CUDA DLLs and import libraries ship; - _C depends on no CUDA DLL and registers no CUDA backend; - no shipped CUDA DLL imports cudart64_*.dll (only the driver), and no Windows NVIDIA requirement is declared; - every DLL with device code carries exactly the row's GPUs, both directions, with the newest also as PTX; the expected list comes from cuda_arch_list.sh keyed by the wheel's +cuXYZ tag, not from the TORCH_CUDA_ARCH_LIST the build consumed; - a C++ application linking executorch::backend_cuda gets the three DLLs copied beside it and sees CudaBackend registered; - a program lowered on Linux for Windows runs through the C++ SDK and matches eager, after checking it was exported with the same CUDA train as the wheel (every CUDA 13 train loads cudart64_13.dll, so a mismatch would not fail on its own); - then what a CPU Windows row checks: test_shared_libraries.py (with the three CUDA ownership rows), test_cpp_sdk.py and the MobileNetV3/XNNPACK model run. The Arm Cortex-M Python module is not built with Ninja, so its install check is left out. ## Test plan On Windows 11 x64 with an RTX 5080 (sm_120, driver 595.95), VS 2022 BuildTools, CPython 3.12, and WSL Ubuntu for the export, once per CUDA train: | train | Windows nvcc | WSL export torch | wheel | test_cuda_windows.py | Linux-exported program | |---|---|---|---|---|---| | 13.0 | 13.0.88 (installer) | 2.14.1+cu130 | executorch-1.6.0+cu130 | 52 checks pass | runs on the GPU, maxdiff 2.4e-7 | | 13.2 | 13.2.86 (redist) | 2.14.1+cu132 | executorch-1.6.0+cu132 | 52 checks pass | runs on the GPU, maxdiff 2.4e-7 | | 13.4 | 13.4.59 (redist) | 2.14.0.dev20260810+cu134 | executorch-1.6.0+cu134 | 52 checks pass | runs on the GPU, maxdiff 2.4e-7 | - Each wheel built with TORCH_CUDA_ARCH_LIST from cuda_arch_list.sh: device code for exactly sm_80/86/89/90/100/120, plus sm_120 PTX. 13.4 runs on a driver that reports 13.2, through CUDA's minor-version compatibility. - Before the /Zc:preprocessor change, 13.2 and 13.4 failed to compile the .cu shims with CCCL's "traditional preprocessor" #error; 13.0 was unaffected. - Negative: the cu134 wheel pointed at the 13.0-exported program fails the train check. - The export used `test_cuda_windows.py --export` from this branch in WSL, with mingw-w64 and the Windows CUDA runtime of the same train (13.0 from install_cuda_windows_cross_compile.sh, 13.2/13.4 from the redist archives). - The new CI end-to-end jobs could not be run before upload; they follow cuda-windows.yml's linux_job_v3 / windows_job pairing and the steps above. - Linux: the CMake changes are behind WIN32 / MSVC, or pick the same suffix there (empty on a Linux host); setup.py's new path runs only on Windows.
Stacked on #23121, which ships linkable libraries in the Windows wheel. A Windows user can run a CUDA program but cannot lower one: lowering compiles the model with a toolchain only the Linux side has, and the CUDA Windows CI job already exports on Linux and runs on Windows for exactly this reason. What was missing is a wheel that gives that Windows user the runtime. This adds a Windows CUDA wheel carrying the CUDA delegate as C++ libraries, the same components the Linux CUDA wheel ships (#21645, #21668), and nothing CUDA in the Python module, which on Windows has no CUDA program to lower. The export (AOT) side is unchanged and identical to the CPU wheel's; only the runtime links CUDA. lib/executorch_backend_cuda.dll + .lib the delegate executorch::backend_cuda lib/executorch_extension_cuda.dll + .lib the stream helper executorch::extension_cuda backends/cuda/aoti_cuda_shims.dll the AOTI shim layer (runtime DLL of the delegate) data/lib/aoti_cuda_shims.lib already shipped as the lowering stub; also the shim layer's import library here extension/pybindings/_C.pyd unchanged: no CUDA dependency, no CudaBackend A C++ application links it exactly as on Linux: find_package(executorch REQUIRED COMPONENTS kernels_optimized backend_cuda) target_link_libraries(app PRIVATE executorch::runtime executorch::kernels_optimized executorch::backend_cuda) # plus the $<TARGET_RUNTIME_DLLS:app> copy from the CPU PR, which now also brings the # stream helper and the shim layer The wheel's DLLs import no CUDA runtime DLL: CMake links cudart.lib into them on Windows, and that reaches the driver (nvcuda.dll). The compiled model is different: the library AOTInductor builds into model.pte imports cudart64_13.dll, which comes from the CUDA Toolkit's bin directory on PATH, as the guide says. PyPI's nvidia-cuda-runtime has win_amd64 builds that carry that DLL, but a C++ program does not search site-packages, so declaring it would not make it loadable; the wheel declares no NVIDIA package for Windows. ## How - Generator. The CUDA toolkit's Visual Studio integration fails compiler identification under the ClangCL toolset on some toolkit and Visual Studio pairs (MSB4023 in its targets file), and the MSVC toolset cannot compile the Python extension's sources. A Windows CUDA build therefore uses Ninja Multi-Config with clang-cl, the compiler the CPU wheel uses, and nvcc with cl.exe as its host compiler, when ninja is on PATH. Without it, as in a --no-build-isolation source install, the build keeps the Visual Studio generator it always used. CMAKE_CUDA_COMPILER is passed only when CUDACXX is unset, so options a user put in CUDACXX (such as -allow-unsupported-compiler) still apply. The toolkit is the one install_utils already reports, so what compiles is the train packaging declares. - CUDA 13.2 and newer. Their CCCL stops with an #error under cl.exe's traditional preprocessor, so the CUDA sources pass -Xcompiler=/Zc:preprocessor on Windows (13.0 accepts it too). - flatc / flatcc byproducts and imported locations share one host executable suffix chosen by CMAKE_HOST_WIN32. CMAKE_EXECUTABLE_SUFFIX is the target's: a Zephyr toolchain sets .elf, and naming the byproduct after it left Ninja with no rule for the host tool. - The shims resolve the runtime from executorch.dll rather than a static core. (The PAL's strong symbols are in #23121; C++20 for the delegate and shims comes from #23135.) - The shims keep WINDOWS_EXPORT_ALL_SYMBOLS OFF in a static build, as upstream sets it, and turn it ON in the shared layout: the delegate DLL imports the stream guard functions (get/set/peek/clearCurrentCUDAStream) from the shim DLL, and those carry no AOTI_SHIM_EXPORT, so the annotations alone would not export them. - #23121 keeps a Windows CUDA build static; this drops that guard, so the pybind preset now builds the shared layout with CUDA too, and adds the CUDA import libraries to the Windows packaging list. - extension/cuda links CUDA::cudart before backends/cuda finds the toolkit, so the root finds it first; a CPU torch does not define that target. - The package config adds the stream helper and the shim layer to the delegate's runtime DLL set, and reports EXECUTORCH_RUNTIME_DLLS_EXTRA for the pre-3.28 route, since the shim layer ships outside lib/. ## CI - build-wheels-cuda-windows.yml mirrors the Linux CUDA workflow on the Windows wheel builder and reuses filter_cuda_matrix.py. pre_build_script.sh drops the blanket Windows CUDA=OFF (a CPU row is still forced off by the generic row classification), resolves TORCH_CUDA_ARCH_LIST from cuda_arch_list.sh, and writes it to BUILD_ENV_FILE as well as GITHUB_ENV: the reusable workflow sources that file after the hook, and its own list had replaced the row's (sm_75 in, sm_89 out). It also installs ninja for the generator. - End-to-end, per published CUDA train, the route a Windows user has, modeled on cuda-windows.yml. The trains are read from filter_cuda_matrix.py's SUPPORTED_CUDA_VERSIONS (today cu130, cu132, cu134), the same list the release rows are built from, so adding or dropping a train there changes what is tested. (install_utils.py's list also has 12.6 for source builds, but 12.6 is no longer a published wheel train.) A Linux GPU job exports a Windows-target program with that train's torch (`test_cuda_windows.py --export`), and a Windows GPU job builds the wheel with that train's nvcc, installs it and runs test_cuda_windows.py with the program (test_cuda_windows_e2e.ps1). Trains the images do not carry are assembled from NVIDIA's checksummed redistributable archives by install_cuda_redist.py, which needs no installer or admin rights and fails on a train it has no release for. Each train's Windows job waits only for its own program, so one train failing to export does not skip the others. The wheel build's own smoke test has no GPU program and says so. ## Tests test_cuda_windows.py, the Windows counterpart of test_cuda_linux.py: - the CUDA DLLs and import libraries ship; - _C depends on no CUDA DLL and registers no CUDA backend; - no shipped CUDA DLL imports cudart64_*.dll (only the driver), and no Windows NVIDIA requirement is declared; - every DLL with device code carries exactly the row's GPUs, both directions, with the newest also as PTX; the expected list comes from cuda_arch_list.sh keyed by the wheel's +cuXYZ tag, not from the TORCH_CUDA_ARCH_LIST the build consumed; - a C++ application linking executorch::backend_cuda gets the three DLLs copied beside it and sees CudaBackend registered; - a program lowered on Linux for Windows runs through the C++ SDK and matches eager, after checking it was exported with the same CUDA train as the wheel (every CUDA 13 train loads cudart64_13.dll, so a mismatch would not fail on its own); - then what a CPU Windows row checks: test_shared_libraries.py (with the three CUDA ownership rows), test_cpp_sdk.py and the MobileNetV3/XNNPACK model run. The Arm Cortex-M Python module is not built with Ninja, so its install check is left out. ## Test plan On Windows 11 x64 with an RTX 5080 (sm_120, driver 595.95), VS 2022 BuildTools, CPython 3.12, and WSL Ubuntu for the export, once per CUDA train: | train | Windows nvcc | WSL export torch | wheel | test_cuda_windows.py | Linux-exported program | |---|---|---|---|---|---| | 13.0 | 13.0.88 (installer) | 2.14.1+cu130 | executorch-1.6.0+cu130 | 52 checks pass | runs on the GPU, maxdiff 2.4e-7 | | 13.2 | 13.2.86 (redist) | 2.14.1+cu132 | executorch-1.6.0+cu132 | 52 checks pass | runs on the GPU, maxdiff 2.4e-7 | | 13.4 | 13.4.59 (redist) | 2.14.0.dev20260810+cu134 | executorch-1.6.0+cu134 | 52 checks pass | runs on the GPU, maxdiff 2.4e-7 | - Each wheel built with TORCH_CUDA_ARCH_LIST from cuda_arch_list.sh: device code for exactly sm_80/86/89/90/100/120, plus sm_120 PTX. 13.4 runs on a driver that reports 13.2, through CUDA's minor-version compatibility. - Before the /Zc:preprocessor change, 13.2 and 13.4 failed to compile the .cu shims with CCCL's "traditional preprocessor" #error; 13.0 was unaffected. - Negative: the cu134 wheel pointed at the 13.0-exported program fails the train check. - The export used `test_cuda_windows.py --export` from this branch in WSL, with mingw-w64 and the Windows CUDA runtime of the same train (13.0 from install_cuda_windows_cross_compile.sh, 13.2/13.4 from the redist archives). - The new CI end-to-end jobs could not be run before upload; they follow cuda-windows.yml's linux_job_v3 / windows_job pairing and the steps above. - Linux: the CMake changes are behind WIN32 / MSVC, or pick the same suffix there (empty on a Linux host); setup.py's new path runs only on Windows.
The Windows wheel shipped one fused Python extension and no linkable libraries, so a C++ application got nothing from it: the headers and the CMake package were there, but every component a consumer asked for resolved to nothing. Linux and macOS already ship the runtime, the kernels, the delegates and the thread pool as separate libraries that a C++ application links directly (#21610, #21771). This does the same for Windows. Before (Windows x64, CPython 3.12): extension/pybindings/_C.cp312-win_amd64.pyd 8.27 MB <- everything fused in here lib/ does not exist After: lib/executorch_kernels_optimized.dll 5.02 MB lib/executorch_backend_xnnpack.dll 2.47 MB lib/executorch.dll 0.38 MB lib/executorch_kernels_quantized.dll 0.22 MB lib/executorch_threadpool.dll 0.21 MB lib/executorch_etdump.dll 0.05 MB (loaded by _C; not offered to C++) lib/*.lib <- import libraries a consumer links extension/pybindings/_C.cp312-win_amd64.pyd 0.63 MB <- just the bindings now ## How The three mechanisms that made the shared layout Linux and macOS only now have Windows equivalents: what ELF / Mach-O Windows exporting the runtime API default visibility WINDOWS_EXPORT_ALL_SYMBOLS keeping a registration-only lib --no-as-needed / named dylib /INCLUDE of a per-DLL anchor finding sibling libraries $ORIGIN / @loader_path os.add_dll_directory, and a consumer copies the DLLs - Exports. The runtime has no dllexport annotations, which is why CMakeLists.txt refused a shared Windows build. Each shipped component DLL now exports all of its symbols, unless its target already exports through its own annotations (WINDOWS_EXPORT_ALL_SYMBOLS OFF), which it keeps. CMake builds that list from a target's own objects and cannot read an archive, so the runtime DLL takes its components as objects rather than through /WHOLEARCHIVE. The caller says which archives are part of a DLL's exports (the helper's EXPORTED option) instead of the helper inferring it from a property another call sets, so the result does not depend on call order. That uses $<COMPILE_ONLY>, new in CMake 3.27, so a shared Windows build asks for 3.27 with a clear message, the Windows build requirement is raised to match, and executorch_shared publishes the cxx_std_20 its headers need on Windows. The PAL source is built with its C functions strong there: the export list skips weak symbols, so the DLL exported none of the et_pal_* functions and clock.h's inline ticks_to_ns() failed to link against it. - Retention. A PE import survives only if some symbol from it is referenced, so each shipped DLL exports `executorch_anchor_<name>` and the imported targets carry `/INCLUDE` of it, the counterpart of the Linux `--no-as-needed`. The name follows the shipped file name without any configuration postfix, and is resolved at the end of configure and written out literally, so it survives `cmake --install` (a generator expression would be evaluated against the imported target, which has no OUTPUT_NAME) and is the same in every configuration of a multi-config build. - One registry. A consumer names the runtime's import library ahead of any archive, as on Linux, so nothing resolves the registry from a private static copy. Without this the optimized kernels registered into their own table: 28 operators visible instead of 242. - Loading. A DLL records no search path. The Python entry points that load a DLL depending on executorch/lib register that directory first (portable_lib, kernels.quantized, codegen.tools, llm.custom_ops.op_tile_crop_aot), computed without resolving symlinks so an editable install, where those directories are links, registers its own lib too. A C++ consumer copies the DLLs beside its executable with `$<TARGET_RUNTIME_DLLS>`, which the C++ guide now shows, including in the quick start. - Debug. The DLLs use the C++ library of the configuration the wheel was built in, and a consumer in the other configuration would mix the two and corrupt memory. The package defines ET_PREBUILT_RELEASE_CRT, or ET_PREBUILT_DEBUG_CRT for a DEBUG=1 build, and the runtime headers refuse the mismatched configuration at compile time with a message naming the right one, for both single-config (-DCMAKE_BUILD_TYPE) and multi-config (--config) generators. naming the right one. (A /FAILIFMISMATCH link option would not do: link.exe ignores it on the command line and honours it only inside an object file.) - cpuinfo. The thread pool DLL exports pthreadpool, so the process has one pool, but keeps cpuinfo private. cpuinfo's feature checks (cpuinfo_has_x86_avx2 and the rest) are inline reads of cpuinfo_isa, which cpuinfo.h does not declare dllimport, so a DLL can only read its own copy; with cpuinfo exported, XNNPACK initialized the thread pool's copy and read its own, all zeros, and ran baseline kernels. Each DLL that uses cpuinfo now carries all of it. Measured on an x64 machine with AVX2: an XNNPACK-delegated 2 x Linear(1024) model at batch 256 went from 2.05-2.14 ms to 1.47-1.52 ms per run, with the same results. CUDA is not part of this. A Windows build with CUDA on keeps the static layout it had, since the Windows CUDA delegate has not been built as a DLL; a follow-up stacked on this ships it. Two bugs only the shared layout exposes on Windows are fixed here: - The process hung at exit about half the time. Windows terminates worker threads before running a DLL's static destructors, so destroying the global thread pool there waited on a lock a terminated worker could hold (`LdrShutdownProcess -> pthreadpool_destroy -> mtx_lock`). The pool is now deliberately not destroyed on Windows; Linux and macOS are unchanged. - `quantized_ops_aot_lib.dll` failed to load, which `executorch.kernels.quantized` swallows, so quantized export silently lost its out variants. The Windows wheel is built without the event tracer (unchanged), so the `etdump` component is not offered to C++ there rather than handing a consumer a profiler that records nothing. The delegates are unchanged too: the Windows wheel ships XNNPACK, as it did before this change. QNN, OpenVINO and TorchAO stay Linux (and for TorchAO, aarch64) wheel components, and Core ML and MLX macOS ones. QNN in particular builds on Windows (build-qnn-windows-x64 and -arm64 pass), but the wheel enables it only where pre_build_script.sh downloads the SDK and the pybind preset's Linux branch turns it on; bringing it to the Windows wheel means the SDK download on the Windows builder, its import library and anchor, and tests, which is a separate change. The pkg-config file the Linux and macOS wheels ship (#23169) is not generated on Windows either, since its link flags are GNU-style and a DLL records no search path for them; the test checks that it is absent. The pre-3.28 variables route links the import libraries and states C++20, which the Windows runtime headers need. ## Tests The two wheel suites now run on Windows, as they do on macOS since #21771, wired into test_windows.py. Their Windows forms ask the platform's own questions: - symbols: dumpbin /exports matched against MSVC decorated names, and for code a DLL takes from an archive, which the export list cannot name, the import direction: the owner must import `register_kernels` / `register_backend` from executorch.dll; - dependencies: dumpbin /dependents and /imports in place of readelf / otool; - loading: LoadLibrary per binary in its own process, which binds every import as ldd -r does, from the package and from a relocated copy; - paths: every recorded dependency is a bare DLL name; - platform tag: the PE machine field of every binary against win_amd64; - C++ consumers: built with --config Release, DLLs copied beside the executable, and the installed package removed from PATH so nothing resolves it for them; - package entry points: the extensions are imported the way a user reaches them (the pybindings through portable_lib), with nothing added to the DLL search path, and `import executorch.kernels.quantized` has to register the quantized out variants, so the entry points' own registration is what is tested; - a consumer in the wrong configuration has to be refused; an application using the thread pool has to exit five runs in a row; no DLL may import cpuinfo from another (the split state has no numeric symptom, only slower kernels); every program the suites build and every Python probe or export they start has a timeout, so a hang fails as a named timeout rather than the whole job timing out. Compiler, CMake and binary tool invocations are not bounded. ## Test plan On Windows 11 x64, VS 2022 BuildTools with ClangCL, CPython 3.12, from a wheel built with `python setup.py bdist_wheel` and installed into a clean environment: - `.ci/scripts/wheel/test_windows.py`: passes, including every check in test_shared_libraries.py and test_cpp_sdk.py (single owner of each component, _C contains no component, every DLL loads in place and relocated, custom op registers, platform tag; version request, 121 of 123 headers compile, documented example builds, runtime alone has no kernels, kernels / quantized / XNNPACK consumers match eager PyTorch, delegate fails without its component, pre-3.28 route on CMake 3.24, relocation, one shared registry) and MobileNetV3 through XNNPACK matching eager. - The exit hang: 8 of 20 and 10 of 20 runs of two consumers hung before the thread pool fix, 0 of 20 after. - Linux and macOS: the CMake changes are behind WIN32 / MSVC and the Python changes behind sys.platform == "win32"; relying on their wheel rows for confirmation.
bd4757a to
813c476
Compare
Stacked on #23121, which ships linkable libraries in the Windows wheel. A Windows user can run a CUDA program but cannot lower one: lowering compiles the model with a toolchain only the Linux side has, and the CUDA Windows CI job already exports on Linux and runs on Windows for exactly this reason. What was missing is a wheel that gives that Windows user the runtime. This adds a Windows CUDA wheel carrying the CUDA delegate as C++ libraries, the same components the Linux CUDA wheel ships (#21645, #21668), and nothing CUDA in the Python module, which on Windows has no CUDA program to lower. The export (AOT) side is unchanged and identical to the CPU wheel's; only the runtime links CUDA. lib/executorch_backend_cuda.dll + .lib the delegate executorch::backend_cuda lib/executorch_extension_cuda.dll + .lib the stream helper executorch::extension_cuda backends/cuda/aoti_cuda_shims.dll the AOTI shim layer (runtime DLL of the delegate) data/lib/aoti_cuda_shims.lib already shipped as the lowering stub; also the shim layer's import library here extension/pybindings/_C.pyd unchanged: no CUDA dependency, no CudaBackend A C++ application links it exactly as on Linux: find_package(executorch REQUIRED COMPONENTS kernels_optimized backend_cuda) target_link_libraries(app PRIVATE executorch::runtime executorch::kernels_optimized executorch::backend_cuda) # plus the $<TARGET_RUNTIME_DLLS:app> copy from the CPU PR, which now also brings the # stream helper and the shim layer The wheel's DLLs import no CUDA runtime DLL: CMake links cudart.lib into them on Windows, and that reaches the driver (nvcuda.dll). The compiled model is different: the library AOTInductor builds into model.pte imports cudart64_13.dll, which comes from the CUDA Toolkit's bin directory on PATH, as the guide says. PyPI's nvidia-cuda-runtime has win_amd64 builds that carry that DLL, but a C++ program does not search site-packages, so declaring it would not make it loadable; the wheel declares no NVIDIA package for Windows. ## How - Generator. The CUDA toolkit's Visual Studio integration fails compiler identification under the ClangCL toolset on some toolkit and Visual Studio pairs (MSB4023 in its targets file), and the MSVC toolset cannot compile the Python extension's sources. A Windows CUDA build therefore uses Ninja Multi-Config with clang-cl, the compiler the CPU wheel uses, and nvcc with cl.exe as its host compiler, when ninja is on PATH. Without it, as in a --no-build-isolation source install, the build keeps the Visual Studio generator it always used. CMAKE_CUDA_COMPILER is passed only when CUDACXX is unset, so options a user put in CUDACXX (such as -allow-unsupported-compiler) still apply. The toolkit is the one install_utils already reports, so what compiles is the train packaging declares. - CUDA 13.2 and newer. Their CCCL stops with an #error under cl.exe's traditional preprocessor, so the CUDA sources pass -Xcompiler=/Zc:preprocessor on Windows (13.0 accepts it too). - flatc / flatcc byproducts and imported locations share one host executable suffix chosen by CMAKE_HOST_WIN32. CMAKE_EXECUTABLE_SUFFIX is the target's: a Zephyr toolchain sets .elf, and naming the byproduct after it left Ninja with no rule for the host tool. - The shims resolve the runtime from executorch.dll rather than a static core. (The PAL's strong symbols are in #23121; C++20 for the delegate and shims comes from #23135.) - The shims keep WINDOWS_EXPORT_ALL_SYMBOLS OFF in a static build, as upstream sets it, and turn it ON in the shared layout: the delegate DLL imports the stream guard functions (get/set/peek/clearCurrentCUDAStream) from the shim DLL, and those carry no AOTI_SHIM_EXPORT, so the annotations alone would not export them. - #23121 keeps a Windows CUDA build static; this drops that guard, so the pybind preset now builds the shared layout with CUDA too, and adds the CUDA import libraries to the Windows packaging list. - extension/cuda links CUDA::cudart before backends/cuda finds the toolkit, so the root finds it first; a CPU torch does not define that target. - The package config adds the stream helper and the shim layer to the delegate's runtime DLL set, and reports EXECUTORCH_RUNTIME_DLLS_EXTRA for the pre-3.28 route, since the shim layer ships outside lib/. ## CI - build-wheels-cuda-windows.yml mirrors the Linux CUDA workflow on the Windows wheel builder and reuses filter_cuda_matrix.py. pre_build_script.sh drops the blanket Windows CUDA=OFF (a CPU row is still forced off by the generic row classification), resolves TORCH_CUDA_ARCH_LIST from cuda_arch_list.sh, and writes it to BUILD_ENV_FILE as well as GITHUB_ENV: the reusable workflow sources that file after the hook, and its own list had replaced the row's (sm_75 in, sm_89 out). It also installs ninja for the generator. - End-to-end, per published CUDA train, the route a Windows user has, modeled on cuda-windows.yml. The trains are read from filter_cuda_matrix.py's SUPPORTED_CUDA_VERSIONS (today cu132 and cu134; main stopped publishing cu130 in #23380), the same list the release rows are built from, so adding or dropping a train there changes what is tested. (install_utils.py's list also has 12.6 for source builds, but 12.6 is no longer a published wheel train.) A Linux GPU job exports a Windows-target program with that train's torch (`test_cuda_windows.py --export`), and a Windows GPU job builds the wheel with that train's nvcc, installs it and runs test_cuda_windows.py with the program (test_cuda_windows_e2e.ps1). Trains the images do not carry are assembled from NVIDIA's checksummed redistributable archives by install_cuda_redist.py, which needs no installer or admin rights and fails on a train it has no release for. Each train's Windows job waits only for its own program, so one train failing to export does not skip the others. The wheel build's own smoke test has no GPU program and says so. ## Tests test_cuda_windows.py, the Windows counterpart of test_cuda_linux.py: - the CUDA DLLs and import libraries ship; - _C depends on no CUDA DLL and registers no CUDA backend; - no shipped CUDA DLL imports cudart64_*.dll (only the driver), and no Windows NVIDIA requirement is declared; - every DLL with device code carries exactly the row's GPUs, both directions, with the newest also as PTX; the expected list comes from cuda_arch_list.sh keyed by the wheel's +cuXYZ tag, not from the TORCH_CUDA_ARCH_LIST the build consumed; - a C++ application linking executorch::backend_cuda gets the three DLLs copied beside it and sees CudaBackend registered; - a program lowered on Linux for Windows runs through the C++ SDK and matches eager, after checking it was exported with the same CUDA train as the wheel (every CUDA 13 train loads cudart64_13.dll, so a mismatch would not fail on its own); - then what a CPU Windows row checks: test_shared_libraries.py (with the three CUDA ownership rows), test_cpp_sdk.py and the MobileNetV3/XNNPACK model run. The Arm Cortex-M Python module is not built with Ninja, so its install check is left out. ## Test plan On Windows 11 x64 with an RTX 5080 (sm_120, driver 595.95), VS 2022 BuildTools, CPython 3.12, and WSL Ubuntu for the export, once per CUDA train. 13.2 and 13.4 are the published trains and the ones CI covers; 13.0 is no longer published (#23380) but still builds from source, so it was run too: | train | Windows nvcc | WSL export torch | wheel | test_cuda_windows.py | Linux-exported program | |---|---|---|---|---|---| | 13.0 | 13.0.88 (installer) | 2.14.1+cu130 | executorch-1.6.0+cu130 | 52 checks pass | runs on the GPU, maxdiff 2.4e-7 | | 13.2 | 13.2.86 (redist) | 2.14.1+cu132 | executorch-1.6.0+cu132 | 52 checks pass | runs on the GPU, maxdiff 2.4e-7 | | 13.4 | 13.4.59 (redist) | 2.14.0.dev20260810+cu134 | executorch-1.6.0+cu134 | 52 checks pass | runs on the GPU, maxdiff 2.4e-7 | - Each wheel built with TORCH_CUDA_ARCH_LIST from cuda_arch_list.sh: device code for exactly sm_80/86/89/90/100/120, plus sm_120 PTX. 13.4 runs on a driver that reports 13.2, through CUDA's minor-version compatibility. - Before the /Zc:preprocessor change, 13.2 and 13.4 failed to compile the .cu shims with CCCL's "traditional preprocessor" #error; 13.0 was unaffected. - Negative: the cu134 wheel pointed at the 13.0-exported program fails the train check. - The export used `test_cuda_windows.py --export` from this branch in WSL, with mingw-w64 and the Windows CUDA runtime of the same train (13.0 from install_cuda_windows_cross_compile.sh, 13.2/13.4 from the redist archives). - The new CI end-to-end jobs could not be run before upload; they follow cuda-windows.yml's linux_job_v3 / windows_job pairing and the steps above. - Linux: the CMake changes are behind WIN32 / MSVC, or pick the same suffix there (empty on a Linux host); setup.py's new path runs only on Windows.
Stacked on #23121, which ships linkable libraries in the Windows wheel. A Windows user can run a CUDA program but cannot lower one: lowering compiles the model with a toolchain only the Linux side has, and the CUDA Windows CI job already exports on Linux and runs on Windows for exactly this reason. What was missing is a wheel that gives that Windows user the runtime. This adds a Windows CUDA wheel carrying the CUDA delegate as C++ libraries, the same components the Linux CUDA wheel ships (#21645, #21668), and nothing CUDA in the Python module, which on Windows has no CUDA program to lower. The export (AOT) side is unchanged and identical to the CPU wheel's; only the runtime links CUDA. lib/executorch_backend_cuda.dll + .lib the delegate executorch::backend_cuda lib/executorch_extension_cuda.dll + .lib the stream helper executorch::extension_cuda backends/cuda/aoti_cuda_shims.dll the AOTI shim layer (runtime DLL of the delegate) data/lib/aoti_cuda_shims.lib already shipped as the lowering stub; also the shim layer's import library here extension/pybindings/_C.pyd unchanged: no CUDA dependency, no CudaBackend A C++ application links it exactly as on Linux: find_package(executorch REQUIRED COMPONENTS kernels_optimized backend_cuda) target_link_libraries(app PRIVATE executorch::runtime executorch::kernels_optimized executorch::backend_cuda) # plus the $<TARGET_RUNTIME_DLLS:app> copy from the CPU PR, which now also brings the # stream helper and the shim layer The wheel's DLLs import no CUDA runtime DLL: CMake links cudart.lib into them on Windows, and that reaches the driver (nvcuda.dll). The compiled model is different: the library AOTInductor builds into model.pte imports cudart64_13.dll, which comes from the CUDA Toolkit's bin directory on PATH, as the guide says. PyPI's nvidia-cuda-runtime has win_amd64 builds that carry that DLL, but a C++ program does not search site-packages, so declaring it would not make it loadable; the wheel declares no NVIDIA package for Windows. ## How - Generator. The CUDA toolkit's Visual Studio integration fails compiler identification under the ClangCL toolset on some toolkit and Visual Studio pairs (MSB4023 in its targets file), and the MSVC toolset cannot compile the Python extension's sources. A Windows CUDA build therefore uses Ninja Multi-Config with clang-cl, the compiler the CPU wheel uses, and nvcc with cl.exe as its host compiler, when ninja is on PATH. Without it, as in a --no-build-isolation source install, the build keeps the Visual Studio generator it always used. CMAKE_CUDA_COMPILER is passed only when CUDACXX is unset, so options a user put in CUDACXX (such as -allow-unsupported-compiler) still apply. The toolkit is the one install_utils already reports, so what compiles is the train packaging declares. - CUDA 13.2 and newer. Their CCCL stops with an #error under cl.exe's traditional preprocessor, so the CUDA sources pass -Xcompiler=/Zc:preprocessor on Windows (13.0 accepts it too). - flatc / flatcc byproducts and imported locations share one host executable suffix chosen by CMAKE_HOST_WIN32. CMAKE_EXECUTABLE_SUFFIX is the target's: a Zephyr toolchain sets .elf, and naming the byproduct after it left Ninja with no rule for the host tool. - The shims resolve the runtime from executorch.dll rather than a static core. (The PAL's strong symbols are in #23121; C++20 for the delegate and shims comes from #23135.) - The shims keep WINDOWS_EXPORT_ALL_SYMBOLS OFF in a static build, as upstream sets it, and turn it ON in the shared layout: the delegate DLL imports the stream guard functions (get/set/peek/clearCurrentCUDAStream) from the shim DLL, and those carry no AOTI_SHIM_EXPORT, so the annotations alone would not export them. - #23121 keeps a Windows CUDA build static; this drops that guard, so the pybind preset now builds the shared layout with CUDA too, and adds the CUDA import libraries to the Windows packaging list. - extension/cuda links CUDA::cudart before backends/cuda finds the toolkit, so the root finds it first; a CPU torch does not define that target. - The package config adds the stream helper and the shim layer to the delegate's runtime DLL set, and reports EXECUTORCH_RUNTIME_DLLS_EXTRA for the pre-3.28 route, since the shim layer ships outside lib/. ## CI - build-wheels-cuda-windows.yml mirrors the Linux CUDA workflow on the Windows wheel builder and reuses filter_cuda_matrix.py. pre_build_script.sh drops the blanket Windows CUDA=OFF (a CPU row is still forced off by the generic row classification), resolves TORCH_CUDA_ARCH_LIST from cuda_arch_list.sh, and writes it to BUILD_ENV_FILE as well as GITHUB_ENV: the reusable workflow sources that file after the hook, and its own list had replaced the row's (sm_75 in, sm_89 out). It also installs ninja for the generator. - End-to-end, per published CUDA train, the route a Windows user has, modeled on cuda-windows.yml. The trains are read from filter_cuda_matrix.py's SUPPORTED_CUDA_VERSIONS (today cu132 and cu134; main stopped publishing cu130 in #23380), the same list the release rows are built from, so adding or dropping a train there changes what is tested. (install_utils.py's list also has 12.6 for source builds, but 12.6 is no longer a published wheel train.) A Linux GPU job exports a Windows-target program with that train's torch (`test_cuda_windows.py --export`), and a Windows GPU job builds the wheel with that train's nvcc, installs it and runs test_cuda_windows.py with the program (test_cuda_windows_e2e.ps1). Trains the images do not carry are assembled from NVIDIA's checksummed redistributable archives by install_cuda_redist.py, which needs no installer or admin rights and fails on a train it has no release for. Each train's Windows job waits only for its own program, so one train failing to export does not skip the others. The wheel build's own smoke test has no GPU program and says so. ## Tests test_cuda_windows.py, the Windows counterpart of test_cuda_linux.py: - the CUDA DLLs and import libraries ship; - _C depends on no CUDA DLL and registers no CUDA backend; - no shipped CUDA DLL imports cudart64_*.dll (only the driver), and no Windows NVIDIA requirement is declared; - every DLL with device code carries exactly the row's GPUs, both directions, with the newest also as PTX; the expected list comes from cuda_arch_list.sh keyed by the wheel's +cuXYZ tag, not from the TORCH_CUDA_ARCH_LIST the build consumed; - a C++ application linking executorch::backend_cuda gets the three DLLs copied beside it and sees CudaBackend registered; - a program lowered on Linux for Windows runs through the C++ SDK and matches eager, after checking it was exported with the same CUDA train as the wheel (every CUDA 13 train loads cudart64_13.dll, so a mismatch would not fail on its own); - then what a CPU Windows row checks: test_shared_libraries.py (with the three CUDA ownership rows), test_cpp_sdk.py and the MobileNetV3/XNNPACK model run. The Arm Cortex-M Python module is not built with Ninja, so its install check is left out. ## Test plan On Windows 11 x64 with an RTX 5080 (sm_120, driver 595.95), VS 2022 BuildTools, CPython 3.12, and WSL Ubuntu for the export, once per CUDA train. 13.2 and 13.4 are the published trains and the ones CI covers; 13.0 is no longer published (#23380) but still builds from source, so it was run too: | train | Windows nvcc | WSL export torch | wheel | test_cuda_windows.py | Linux-exported program | |---|---|---|---|---|---| | 13.0 | 13.0.88 (installer) | 2.14.1+cu130 | executorch-1.6.0+cu130 | 52 checks pass | runs on the GPU, maxdiff 2.4e-7 | | 13.2 | 13.2.86 (redist) | 2.14.1+cu132 | executorch-1.6.0+cu132 | 52 checks pass | runs on the GPU, maxdiff 2.4e-7 | | 13.4 | 13.4.59 (redist) | 2.14.0.dev20260810+cu134 | executorch-1.6.0+cu134 | 52 checks pass | runs on the GPU, maxdiff 2.4e-7 | - Each wheel built with TORCH_CUDA_ARCH_LIST from cuda_arch_list.sh: device code for exactly sm_80/86/89/90/100/120, plus sm_120 PTX. 13.4 runs on a driver that reports 13.2, through CUDA's minor-version compatibility. - Before the /Zc:preprocessor change, 13.2 and 13.4 failed to compile the .cu shims with CCCL's "traditional preprocessor" #error; 13.0 was unaffected. - Negative: the cu134 wheel pointed at the 13.0-exported program fails the train check. - The export used `test_cuda_windows.py --export` from this branch in WSL, with mingw-w64 and the Windows CUDA runtime of the same train (13.0 from install_cuda_windows_cross_compile.sh, 13.2/13.4 from the redist archives). - The new CI end-to-end jobs could not be run before upload; they follow cuda-windows.yml's linux_job_v3 / windows_job pairing and the steps above. - Linux: the CMake changes are behind WIN32 / MSVC, or pick the same suffix there (empty on a Linux host); setup.py's new path runs only on Windows.
Stacked on #23121, which ships linkable libraries in the Windows wheel. A Windows user can run a CUDA program but cannot lower one: lowering compiles the model with a toolchain only the Linux side has, and the CUDA Windows CI job already exports on Linux and runs on Windows for exactly this reason. What was missing is a wheel that gives that Windows user the runtime. This adds a Windows CUDA wheel carrying the CUDA delegate as C++ libraries, the same components the Linux CUDA wheel ships (#21645, #21668), and nothing CUDA in the Python module, which on Windows has no CUDA program to lower. The export (AOT) side is unchanged and identical to the CPU wheel's; only the runtime links CUDA. lib/executorch_backend_cuda.dll + .lib the delegate executorch::backend_cuda lib/executorch_extension_cuda.dll + .lib the stream helper executorch::extension_cuda backends/cuda/aoti_cuda_shims.dll the AOTI shim layer (runtime DLL of the delegate) data/lib/aoti_cuda_shims.lib already shipped as the lowering stub; also the shim layer's import library here extension/pybindings/_C.pyd unchanged: no CUDA dependency, no CudaBackend A C++ application links it exactly as on Linux: find_package(executorch REQUIRED COMPONENTS kernels_optimized backend_cuda) target_link_libraries(app PRIVATE executorch::runtime executorch::kernels_optimized executorch::backend_cuda) # plus the $<TARGET_RUNTIME_DLLS:app> copy from the CPU PR, which now also brings the # stream helper and the shim layer The wheel's DLLs import no CUDA runtime DLL: CMake links cudart.lib into them on Windows, and that reaches the driver (nvcuda.dll). The compiled model is different: the library AOTInductor builds into model.pte imports cudart64_13.dll, which comes from the CUDA Toolkit's bin directory on PATH, as the guide says. PyPI's nvidia-cuda-runtime has win_amd64 builds that carry that DLL, but a C++ program does not search site-packages, so declaring it would not make it loadable; the wheel declares no NVIDIA package for Windows. ## How - Generator. The CUDA toolkit's Visual Studio integration fails compiler identification under the ClangCL toolset on some toolkit and Visual Studio pairs (MSB4023 in its targets file), and the MSVC toolset cannot compile the Python extension's sources. A Windows CUDA build therefore uses Ninja Multi-Config with clang-cl, the compiler the CPU wheel uses, and nvcc with cl.exe as its host compiler, when ninja is on PATH. Without it, as in a --no-build-isolation source install, the build keeps the Visual Studio generator it always used. CMAKE_CUDA_COMPILER is passed only when CUDACXX is unset, so options a user put in CUDACXX (such as -allow-unsupported-compiler) still apply. The toolkit is the one install_utils already reports, so what compiles is the train packaging declares. - CUDA 13.2 and newer. Their CCCL stops with an #error under cl.exe's traditional preprocessor, so the CUDA sources pass -Xcompiler=/Zc:preprocessor on Windows (13.0 accepts it too). - flatc / flatcc byproducts and imported locations share one host executable suffix chosen by CMAKE_HOST_WIN32. CMAKE_EXECUTABLE_SUFFIX is the target's: a Zephyr toolchain sets .elf, and naming the byproduct after it left Ninja with no rule for the host tool. - The shims resolve the runtime from executorch.dll rather than a static core. (The PAL's strong symbols are in #23121; C++20 for the delegate and shims comes from #23135.) - The shims keep WINDOWS_EXPORT_ALL_SYMBOLS OFF in a static build, as upstream sets it, and turn it ON in the shared layout: the delegate DLL imports the stream guard functions (get/set/peek/clearCurrentCUDAStream) from the shim DLL, and those carry no AOTI_SHIM_EXPORT, so the annotations alone would not export them. - #23121 keeps a Windows CUDA build static; this drops that guard, so the pybind preset now builds the shared layout with CUDA too, and adds the CUDA import libraries to the Windows packaging list. - extension/cuda links CUDA::cudart before backends/cuda finds the toolkit, so the root finds it first; a CPU torch does not define that target. - The package config adds the stream helper and the shim layer to the delegate's runtime DLL set, and reports EXECUTORCH_RUNTIME_DLLS_EXTRA for the pre-3.28 route, since the shim layer ships outside lib/. ## CI - build-wheels-cuda-windows.yml mirrors the Linux CUDA workflow on the Windows wheel builder and reuses filter_cuda_matrix.py. pre_build_script.sh drops the blanket Windows CUDA=OFF (a CPU row is still forced off by the generic row classification), resolves TORCH_CUDA_ARCH_LIST from cuda_arch_list.sh, and writes it to BUILD_ENV_FILE as well as GITHUB_ENV: the reusable workflow sources that file after the hook, and its own list had replaced the row's (sm_75 in, sm_89 out). It also installs ninja for the generator. - End-to-end, per published CUDA train, the route a Windows user has, modeled on cuda-windows.yml. The trains are read from filter_cuda_matrix.py's SUPPORTED_CUDA_VERSIONS (today cu132 and cu134; main stopped publishing cu130 in #23380), the same list the release rows are built from, so adding or dropping a train there changes what is tested. (install_utils.py's list also has 12.6 for source builds, but 12.6 is no longer a published wheel train.) A Linux GPU job exports a Windows-target program with that train's torch (`test_cuda_windows.py --export`), and a Windows GPU job builds the wheel with that train's nvcc, installs it and runs test_cuda_windows.py with the program (test_cuda_windows_e2e.ps1). Trains the images do not carry are assembled from NVIDIA's checksummed redistributable archives by install_cuda_redist.py, which needs no installer or admin rights and fails on a train it has no release for. Each train's Windows job waits only for its own program, so one train failing to export does not skip the others. The wheel build's own smoke test has no GPU program and says so. ## Tests test_cuda_windows.py, the Windows counterpart of test_cuda_linux.py: - the CUDA DLLs and import libraries ship; - _C depends on no CUDA DLL and registers no CUDA backend; - no shipped CUDA DLL imports cudart64_*.dll (only the driver), and no Windows NVIDIA requirement is declared; - every DLL with device code carries exactly the row's GPUs, both directions, with the newest also as PTX; the expected list comes from cuda_arch_list.sh keyed by the wheel's +cuXYZ tag, not from the TORCH_CUDA_ARCH_LIST the build consumed; - a C++ application linking executorch::backend_cuda gets the three DLLs copied beside it and sees CudaBackend registered; - a program lowered on Linux for Windows runs through the C++ SDK and matches eager, after checking it was exported with the same CUDA train as the wheel (every CUDA 13 train loads cudart64_13.dll, so a mismatch would not fail on its own); - then what a CPU Windows row checks: test_shared_libraries.py (with the three CUDA ownership rows), test_cpp_sdk.py and the MobileNetV3/XNNPACK model run. The Arm Cortex-M Python module is not built with Ninja, so its install check is left out. ## Test plan On Windows 11 x64 with an RTX 5080 (sm_120, driver 595.95), VS 2022 BuildTools, CPython 3.12, and WSL Ubuntu for the export, once per CUDA train. 13.2 and 13.4 are the published trains and the ones CI covers; 13.0 is no longer published (#23380) but still builds from source, so it was run too: | train | Windows nvcc | WSL export torch | wheel | test_cuda_windows.py | Linux-exported program | |---|---|---|---|---|---| | 13.0 | 13.0.88 (installer) | 2.14.1+cu130 | executorch-1.6.0+cu130 | 52 checks pass | runs on the GPU, maxdiff 2.4e-7 | | 13.2 | 13.2.86 (redist) | 2.14.1+cu132 | executorch-1.6.0+cu132 | 52 checks pass | runs on the GPU, maxdiff 2.4e-7 | | 13.4 | 13.4.59 (redist) | 2.14.0.dev20260810+cu134 | executorch-1.6.0+cu134 | 52 checks pass | runs on the GPU, maxdiff 2.4e-7 | - Each wheel built with TORCH_CUDA_ARCH_LIST from cuda_arch_list.sh: device code for exactly sm_80/86/89/90/100/120, plus sm_120 PTX. 13.4 runs on a driver that reports 13.2, through CUDA's minor-version compatibility. - Before the /Zc:preprocessor change, 13.2 and 13.4 failed to compile the .cu shims with CCCL's "traditional preprocessor" #error; 13.0 was unaffected. - Negative: the cu134 wheel pointed at the 13.0-exported program fails the train check. - The export used `test_cuda_windows.py --export` from this branch in WSL, with mingw-w64 and the Windows CUDA runtime of the same train (13.0 from install_cuda_windows_cross_compile.sh, 13.2/13.4 from the redist archives). - The new CI end-to-end jobs could not be run before upload; they follow cuda-windows.yml's linux_job_v3 / windows_job pairing and the steps above. - Linux: the CMake changes are behind WIN32 / MSVC, or pick the same suffix there (empty on a Linux host); setup.py's new path runs only on Windows.
shoumikhin
left a comment
There was a problem hiding this comment.
This looks good to me. Two small things inline, neither blocks.
| set_target_properties( | ||
| ${target_name} PROPERTIES EXECUTORCH_ANCHOR "${_anchor}" | ||
| ) | ||
| target_link_options(${target_name} INTERFACE "LINKER:/INCLUDE:${_anchor}") |
There was a problem hiding this comment.
minor
I have not run this on Windows, but I expect import executorch.extension.pybindings.data_loader as the first import to fail with a DLL load error now. That module uses no runtime function, yet it gets this /INCLUDE, so it imports executorch_anchor_executorch from executorch.dll. Nothing on its import path adds the lib directory to the DLL search path: its packages have no __init__.py, and it is not one of the four entry points your description lists. The import test imports portable_lib first, so it does not see this. Skipping the anchor for this module, or adding the directory for it, should fix it.
| // program built with the other one (/MDd and /MTd define _DEBUG, /MD and /MT do | ||
| // not) sees differently laid out types: the two mixed in one process corrupt | ||
| // memory instead of failing to link. | ||
| #if defined(ET_PREBUILT_RELEASE_CRT) && defined(_DEBUG) |
There was a problem hiding this comment.
minor
This checks _DEBUG only. A Release program built with /D_ITERATOR_DEBUG_LEVEL=1 still compiles against the package, and at that level the standard containers carry an extra pointer, so their layout differs from the DLLs'. Checking _ITERATOR_DEBUG_LEVEL too would catch it.
Ship linkable libraries in the Windows wheel
The Windows wheel shipped one fused Python extension and no linkable libraries, so a C++
application got nothing from it: the headers and the CMake package were there, but every
component a consumer asked for resolved to nothing. Linux and macOS already ship the runtime,
the kernels, the delegates and the thread pool as separate libraries that a C++ application
links directly (#21610, #21771). This does the same for Windows.
Before (Windows x64, CPython 3.12):
After:
How
The three mechanisms that made the shared layout Linux and macOS only now have Windows
equivalents:
shared Windows build. Each shipped component DLL now exports all of its symbols, unless
its target already exports through its own annotations (WINDOWS_EXPORT_ALL_SYMBOLS OFF),
which it keeps. CMake builds that list from a target's own objects and cannot read an
archive, so the runtime DLL takes its components as objects rather than through
/WHOLEARCHIVE. The caller says which archives are part of a DLL's exports (the helper's
EXPORTED option) instead of the helper inferring it from a property another call sets,
so the result does not depend on call order. That uses $<COMPILE_ONLY>, new in CMake
3.27, so a shared Windows build asks for 3.27 with a clear message, the Windows build
requirement is raised to match, and executorch_shared publishes the cxx_std_20 its
headers need on Windows. The PAL source is built with its C functions strong there:
the export list skips weak symbols, so the DLL exported none of the et_pal_* functions
and clock.h's inline ticks_to_ns() failed to link against it.
shipped DLL exports
executorch_anchor_<name>and the imported targets carry/INCLUDEof it, the counterpart of the Linux--no-as-needed. The name follows theshipped file name without any configuration postfix, and is resolved at the end of
configure and written out literally, so it survives
cmake --install(a generatorexpression would be evaluated against the imported target, which has no OUTPUT_NAME)
and is the same in every configuration of a multi-config build.
Linux, so nothing resolves the registry from a private static copy. Without this the
optimized kernels registered into their own table: 28 operators visible instead of 242.
on executorch/lib register that directory first (portable_lib, kernels.quantized,
codegen.tools, llm.custom_ops.op_tile_crop_aot), computed without resolving symlinks so
an editable install, where those directories are links, registers its own lib too. A C++
consumer copies the DLLs beside its executable with
$<TARGET_RUNTIME_DLLS>, which theC++ guide now shows, including in the quick start.
consumer in the other configuration would mix the two and corrupt memory. The package
defines ET_PREBUILT_RELEASE_CRT, or ET_PREBUILT_DEBUG_CRT for a DEBUG=1 build, and the
runtime headers refuse the mismatched configuration at compile time with a message
naming the right one, for both single-config (-DCMAKE_BUILD_TYPE) and multi-config
(--config) generators.
naming the right one. (A /FAILIFMISMATCH link option would not do: link.exe ignores it
on the command line and honours it only inside an object file.)
cpuinfo private. cpuinfo's feature checks (cpuinfo_has_x86_avx2 and the rest) are
inline reads of cpuinfo_isa, which cpuinfo.h does not declare dllimport, so a DLL can
only read its own copy; with cpuinfo exported, XNNPACK initialized the thread pool's copy
and read its own, all zeros, and ran baseline kernels. Each DLL that uses cpuinfo now
carries all of it. Measured on an x64 machine with AVX2: an XNNPACK-delegated
2 x Linear(1024) model at batch 256 went from 2.05-2.14 ms to 1.47-1.52 ms per run, with
the same results.
CUDA is not part of this. A Windows build with CUDA on keeps the static layout it had,
since the Windows CUDA delegate has not been built as a DLL; a follow-up stacked on this
ships it.
Two bugs only the shared layout exposes on Windows are fixed here:
running a DLL's static destructors, so destroying the global thread pool there waited on a
lock a terminated worker could hold (
LdrShutdownProcess -> pthreadpool_destroy -> mtx_lock). The pool is now deliberately not destroyed on Windows; Linux and macOS areunchanged.
quantized_ops_aot_lib.dllfailed to load, whichexecutorch.kernels.quantizedswallows,so quantized export silently lost its out variants.
The Windows wheel is built without the event tracer (unchanged), so the
etdumpcomponentis not offered to C++ there rather than handing a consumer a profiler that records nothing.
The delegates are unchanged too: the Windows wheel ships XNNPACK, as it did before this
change. QNN, OpenVINO and TorchAO stay Linux (and for TorchAO, aarch64) wheel components, and
Core ML and MLX macOS ones. QNN in particular builds on Windows (build-qnn-windows-x64 and
-arm64 pass), but the wheel enables it only where pre_build_script.sh downloads the SDK and
the pybind preset's Linux branch turns it on; bringing it to the Windows wheel means the SDK
download on the Windows builder, its import library and anchor, and tests, which is a
separate change.
The pkg-config file the Linux and macOS wheels ship (#23169) is not generated on Windows
either, since its link flags are GNU-style and a DLL records no search path for them; the
test checks that it is absent.
The pre-3.28 variables route links the import libraries and states C++20, which the Windows
runtime headers need.
Tests
The two wheel suites now run on Windows, as they do on macOS since #21771, wired into
test_windows.py. Their Windows forms ask the platform's own questions:
from an archive, which the export list cannot name, the import direction: the owner must
import
register_kernels/register_backendfrom executorch.dll;does, from the package and from a relocated copy;
installed package removed from PATH so nothing resolves it for them;
pybindings through portable_lib), with nothing added to the DLL search path, and
import executorch.kernels.quantizedhas to register the quantized out variants, so theentry points' own registration is what is tested;
thread pool has to exit five runs in a row; no DLL may import cpuinfo from another
(the split state has no numeric symptom, only slower kernels); every program the suites
build and every Python probe or export they start has a timeout, so a hang fails as a
named timeout rather than the whole job timing out. Compiler, CMake and binary tool
invocations are not bounded.
Test plan
On Windows 11 x64, VS 2022 BuildTools with ClangCL, CPython 3.12, from a wheel built with
python setup.py bdist_wheeland installed into a clean environment:.ci/scripts/wheel/test_windows.py: passes, including every check intest_shared_libraries.py and test_cpp_sdk.py (single owner of each component, _C contains
no component, every DLL loads in place and relocated, custom op registers, platform tag;
version request, 121 of 123 headers compile, documented example builds, runtime alone has
no kernels, kernels / quantized / XNNPACK consumers match eager PyTorch, delegate fails
without its component, pre-3.28 route on CMake 3.24, relocation, one shared registry) and
MobileNetV3 through XNNPACK matching eager.
0 of 20 after.
sys.platform == "win32"; relying on their wheel rows for confirmation.