Skip to content

Ship the CUDA runtime in the Windows CUDA wheel - #23123

Open
Gasoonjia wants to merge 1 commit into
gasoonjia/windows-cpu-sdk-wheelfrom
gasoonjia/windows-cuda-sdk-wheel
Open

Gasoonjia wants to merge 1 commit into
gasoonjia/windows-cpu-sdk-wheelfrom
gasoonjia/windows-cuda-sdk-wheel

Conversation

@Gasoonjia

@Gasoonjia Gasoonjia commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

Stacked on #23121, which ships linkable libraries in the Windows wheel.

A Windows user can run a CUDA program but cannot lower one: lowering compiles the model
with a toolchain only the Linux side has, and the CUDA Windows CI job already exports on
Linux and runs on Windows for exactly this reason. What was missing is a wheel that gives
that Windows user the runtime. This adds a Windows CUDA wheel carrying the CUDA delegate
as C++ libraries, the same components the Linux CUDA wheel ships (#21645, #21668), and
nothing CUDA in the Python module, which on Windows has no CUDA program to lower. The
export (AOT) side is unchanged and identical to the CPU wheel's; only the runtime links CUDA.

lib/executorch_backend_cuda.dll    + .lib   the delegate          executorch::backend_cuda
lib/executorch_extension_cuda.dll  + .lib   the stream helper     executorch::extension_cuda
backends/cuda/aoti_cuda_shims.dll            the AOTI shim layer   (runtime DLL of the delegate)
data/lib/aoti_cuda_shims.lib                 already shipped as the lowering stub; also the
                                             shim layer's import library here
extension/pybindings/_C.pyd                  unchanged: no CUDA dependency, no CudaBackend

A C++ application links it exactly as on Linux:

find_package(executorch REQUIRED COMPONENTS kernels_optimized backend_cuda)
target_link_libraries(app PRIVATE executorch::runtime executorch::kernels_optimized
                                  executorch::backend_cuda)
# plus the $<TARGET_RUNTIME_DLLS:app> copy from the CPU PR, which now also brings the
# stream helper and the shim layer

The wheel's DLLs import no CUDA runtime DLL: CMake links cudart.lib into them on Windows,
and that reaches the driver (nvcuda.dll). The compiled model is different: the library
AOTInductor builds into model.pte imports cudart64_13.dll, which comes from the CUDA
Toolkit's bin directory on PATH, as the guide says. PyPI's nvidia-cuda-runtime has win_amd64
builds that carry that DLL, but a C++ program does not search site-packages, so declaring it
would not make it loadable; the wheel declares no NVIDIA package for Windows.

How

  • Generator. The CUDA toolkit's Visual Studio integration fails compiler identification
    under the ClangCL toolset on some toolkit and Visual Studio pairs (MSB4023 in its targets
    file), and the MSVC toolset cannot compile the Python extension's sources. A Windows CUDA
    build therefore uses Ninja Multi-Config with clang-cl, the compiler the CPU wheel uses,
    and nvcc with cl.exe as its host compiler, when ninja is on PATH. Without it, as in a
    --no-build-isolation source install, the build keeps the Visual Studio generator it always
    used. CMAKE_CUDA_COMPILER is passed only when CUDACXX is unset, so options a user put in
    CUDACXX (such as -allow-unsupported-compiler) still apply. The toolkit is the one
    install_utils already reports, so what compiles is the train packaging declares.
  • CUDA 13.2 and newer. Their CCCL stops with an #error under cl.exe's traditional
    preprocessor, so the CUDA sources pass -Xcompiler=/Zc:preprocessor on Windows (13.0
    accepts it too).
  • flatc / flatcc byproducts and imported locations share one host executable suffix chosen
    by CMAKE_HOST_WIN32. CMAKE_EXECUTABLE_SUFFIX is the target's: a Zephyr toolchain sets
    .elf, and naming the byproduct after it left Ninja with no rule for the host tool.
  • The shims resolve the runtime from executorch.dll rather than a static core. (The PAL's
    strong symbols are in Ship linkable libraries in the Windows CPU wheel #23121; C++20 for the delegate and shims comes from Fix CUDA backend Windows build under C++20 #23135.)
  • The shims keep WINDOWS_EXPORT_ALL_SYMBOLS OFF in a static build, as upstream sets it,
    and turn it ON in the shared layout: the delegate DLL imports the stream guard functions
    (get/set/peek/clearCurrentCUDAStream) from the shim DLL, and those carry no
    AOTI_SHIM_EXPORT, so the annotations alone would not export them.
  • Ship linkable libraries in the Windows CPU wheel #23121 keeps a Windows CUDA build static; this drops that guard, so the pybind preset
    now builds the shared layout with CUDA too, and adds the CUDA import libraries to the
    Windows packaging list.
  • extension/cuda links CUDA::cudart before backends/cuda finds the toolkit, so the root
    finds it first; a CPU torch does not define that target.
  • The package config adds the stream helper and the shim layer to the delegate's runtime
    DLL set, and reports EXECUTORCH_RUNTIME_DLLS_EXTRA for the pre-3.28 route, since the
    shim layer ships outside lib/.

CI

  • build-wheels-cuda-windows.yml mirrors the Linux CUDA workflow on the Windows wheel
    builder and reuses filter_cuda_matrix.py. pre_build_script.sh drops the blanket Windows
    CUDA=OFF (a CPU row is still forced off by the generic row classification), resolves
    TORCH_CUDA_ARCH_LIST from cuda_arch_list.sh, and writes it to BUILD_ENV_FILE as well as
    GITHUB_ENV: the reusable workflow sources that file after the hook, and its own list had
    replaced the row's (sm_75 in, sm_89 out). It also installs ninja for the generator.
  • End-to-end, per published CUDA train, the route a Windows user has, modeled on
    cuda-windows.yml. The trains are read from filter_cuda_matrix.py's SUPPORTED_CUDA_VERSIONS
    (today cu132 and cu134; main stopped publishing cu130 in Stop publishing CUDA 13.0 wheels #23380), the same list the release
    rows are built from, so adding or
    dropping a train there changes what is tested. (install_utils.py's list also has 12.6 for
    source builds, but 12.6 is no longer a published wheel train.) A Linux GPU job exports a
    Windows-target program with that train's torch (test_cuda_windows.py --export), and a
    Windows GPU job builds the wheel with that train's nvcc, installs it and runs
    test_cuda_windows.py with the program (test_cuda_windows_e2e.ps1). Trains the images do not
    carry are assembled from NVIDIA's checksummed redistributable archives by
    install_cuda_redist.py, which needs no installer or admin rights and fails on a train it
    has no release for. Each train's Windows job waits only for its own program, so one train
    failing to export does not skip the others. The Windows GPU runners come without an NVIDIA
    driver (nvcuda.dll is missing and cudaGetDeviceCount returns cudaErrorInsufficientDriver),
    so the job first installs the same data-center driver, from the same place, that
    pytorch/pytorch's Windows CUDA smoke tests install (driver_update.bat, 580.88); cu132 and
    cu134 run on it through CUDA's minor-version compatibility, as PyTorch's own wheels do. The
    wheel build's own smoke test has no GPU program and says so.

Tests

test_cuda_windows.py, the Windows counterpart of test_cuda_linux.py:

  • the CUDA DLLs and import libraries ship;
  • _C depends on no CUDA DLL and registers no CUDA backend;
  • no shipped CUDA DLL imports cudart64_*.dll (only the driver), and no Windows NVIDIA
    requirement is declared;
  • every DLL with device code carries exactly the row's GPUs, both directions, with the
    newest also as PTX; the expected list comes from cuda_arch_list.sh keyed by the wheel's
    +cuXYZ tag, not from the TORCH_CUDA_ARCH_LIST the build consumed;
  • a C++ application linking executorch::backend_cuda gets the three DLLs copied beside it
    and sees CudaBackend registered;
  • a program lowered on Linux for Windows runs through the C++ SDK and matches eager, after
    checking it was exported with the same CUDA train as the wheel (every CUDA 13 train loads
    cudart64_13.dll, so a mismatch would not fail on its own);
  • then what a CPU Windows row checks: test_shared_libraries.py (with the three CUDA
    ownership rows), test_cpp_sdk.py and the MobileNetV3/XNNPACK model run. The Arm Cortex-M
    Python module is not built with Ninja, so its install check is left out.

Test plan

On Windows 11 x64 with an RTX 5080 (sm_120, driver 595.95), VS 2022 BuildTools, CPython 3.12,
and WSL Ubuntu for the export, once per CUDA train. 13.2 and 13.4 are the published trains
and the ones CI covers; 13.0 is no longer published (#23380) but still builds from source, so
it was run too:

train Windows nvcc WSL export torch wheel test_cuda_windows.py Linux-exported program
13.0 13.0.88 (installer) 2.14.1+cu130 executorch-1.6.0+cu130 52 checks pass runs on the GPU, maxdiff 2.4e-7
13.2 13.2.86 (redist) 2.14.1+cu132 executorch-1.6.0+cu132 52 checks pass runs on the GPU, maxdiff 2.4e-7
13.4 13.4.59 (redist) 2.14.0.dev20260810+cu134 executorch-1.6.0+cu134 52 checks pass runs on the GPU, maxdiff 2.4e-7
  • Each wheel built with TORCH_CUDA_ARCH_LIST from cuda_arch_list.sh: device code for exactly
    sm_80/86/89/90/100/120, plus sm_120 PTX. 13.4 runs on a driver that reports 13.2, through
    CUDA's minor-version compatibility.
  • Before the /Zc:preprocessor change, 13.2 and 13.4 failed to compile the .cu shims with
    CCCL's "traditional preprocessor" #error; 13.0 was unaffected.
  • Negative: the cu134 wheel pointed at the 13.0-exported program fails the train check.
  • The export used test_cuda_windows.py --export from this branch in WSL, with mingw-w64 and
    the Windows CUDA runtime of the same train (13.0 from install_cuda_windows_cross_compile.sh,
    13.2/13.4 from the redist archives).
  • The new CI end-to-end jobs could not be run before upload; they follow cuda-windows.yml's
    linux_job_v3 / windows_job pairing and the steps above.
  • Linux: the CMake changes are behind WIN32 / MSVC, or pick the same suffix there (empty on a
    Linux host); setup.py's new path runs only on Windows.

@pytorch-bot

pytorch-bot Bot commented Sep 24, 2026 •

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/23123

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 Pending, 5 Unclassified Failures

As of commit dd388dd with merge base 3794e44 (image):

UNCLASSIFIED FAILURES - DrCI could not classify the following jobs because the workflow did not run on the merge base. The failures may be pre-existing on trunk or introduced by this PR:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-cla meta-cla Bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Sep 24, 2026
Comment thread third-party/CMakeLists.txt Outdated
Comment thread .ci/scripts/wheel/test_cuda_windows.py Outdated
Comment thread .ci/scripts/wheel/test_cuda_windows.py
Comment thread .ci/scripts/wheel/pre_build_script.sh
Comment thread setup.py
Comment thread .github/workflows/build-wheels-cuda-windows.yml
Comment thread setup.py Outdated
@Gasoonjia
Gasoonjia force-pushed the gasoonjia/windows-cpu-sdk-wheel branch from 24b4561 to bd4757a Compare October 2, 2026 20:00
@Gasoonjia
Gasoonjia force-pushed the gasoonjia/windows-cuda-sdk-wheel branch from c2bace9 to a2310ce Compare October 2, 2026 23:58
@Gasoonjia
Gasoonjia force-pushed the gasoonjia/windows-cuda-sdk-wheel branch from a2310ce to 5d6a847 Compare October 3, 2026 01:45
@Gasoonjia
Gasoonjia force-pushed the gasoonjia/windows-cpu-sdk-wheel branch from bd4757a to 813c476 Compare October 3, 2026 02:23
@Gasoonjia
Gasoonjia force-pushed the gasoonjia/windows-cuda-sdk-wheel branch from 5d6a847 to cc32b23 Compare October 3, 2026 02:23
@Gasoonjia
Gasoonjia force-pushed the gasoonjia/windows-cuda-sdk-wheel branch from cc32b23 to 87e3b7a Compare October 3, 2026 02:25
@Gasoonjia
Gasoonjia force-pushed the gasoonjia/windows-cuda-sdk-wheel branch from 87e3b7a to c9578b4 Compare October 3, 2026 03:07

@shoumikhin shoumikhin left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Two blockers are left, both inline: the duplicate allocator in the Windows source install, and the end to end job that has no NVIDIA driver. The earlier threads are resolved.

# carry no AOTI_SHIM_EXPORT, so the annotations alone would not export them;
# export everything, as every other shipped DLL does.
set_target_properties(
aoti_cuda_shims PROPERTIES WINDOWS_EXPORT_ALL_SYMBOLS ON

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

blocker

On Windows with the CUDA toolkit installed, a source install that builds the CUDA backend with the Visual Studio generator no longer links. CUDA is on by default when nvcc is found, and install_executorch uses that generator because nothing installs ninja for it.

runtime/cuda_allocator.cpp is built into aoti_cuda_shims, and on MSVC it is also built into aoti_cuda_backend a few lines below. Now that the Windows build is shared and the shims DLL exports everything, the linker sees CudaAllocator::instance, memcpy_async and release_cached_memory twice and stops with "duplicate symbol".

CI hides it. All 16 test-models-windows and test-model-cuda-windows-e2e jobs hit it, build no wheel, and still end green. The same test-models-windows jobs on #23121 all build the wheel.

The Ninja wheel build links, but the wheel has the same problem in another form: aoti_cuda_shims.dll and executorch_backend_cuda.dll both export the same 13 CudaAllocator symbols, so the process has two allocators.

I think one condition fixes both: add cuda_allocator.cpp to the delegate only for the static layout, so in the shared layout it takes the allocator from the shims DLL, like it already does for the stream functions. I have not built that change.

uses: pytorch/test-infra/.github/workflows/windows_job.yml@main
with:
timeout: 240
runner: windows.g5.4xlarge.nvidia.gpu

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

blocker

The end to end jobs fail on both published trains because this job has no NVIDIA driver. The diagnostic you added shows it: nvcuda.dll present: False, driver supports CUDA 0, and cudaGetDeviceCount -> 35, which is cudaErrorInsufficientDriver. That is why forward failed with cudaGetDevice failed.

windows_job.yml has no driver step. The wheel build job does: it runs driver_update.bat. Could this job install the driver before the test, and print nvidia-smi, so a driver problem is easy to spot next time?

Until then no CI job has run a CUDA program from this wheel, so the GPU claims in the description rest on your local runs.

# Not gated on the export job's result: that is one result for the whole matrix, so a
# single train failing to export skipped every train here. Each train downloads its own
# program, and a train whose export failed fails here on the missing artifact.
if: ${{ !cancelled() && needs.cuda-trains.result == 'success' }}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

major

!cancelled() removes the usual "needed jobs succeeded" check. On a fork pull request the export job is skipped, but this job still runs and then fails to download an artifact that was never made. cuda-windows.yml has a comment about exactly this trap and gates on the export job's result.

Adding the same fork condition here, or needs.export-cuda-windows-wheel-program.result != 'skipped', keeps the "one train failed, run the others" behavior and fixes forks. This pull request is not from a fork, so its own CI cannot show it.

package-name: executorch
name: ${{ matrix.repository }}
uses: pytorch/test-infra/.github/workflows/build_wheels_windows.yml@main
with:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

major

The Windows CUDA wheel build is too close to its time limit. This job sets no timeout, so it gets the reusable workflow's default of 60 minutes. The same job took 58 minutes on an earlier push and hit the 60 minute limit on the latest one, in the smoke test, so no wheel was uploaded. On the latest push the smoke test started 55 minutes in, and it took 9 minutes the time before.

Could you set timeout: 90 here, like the export job does?

```

That export also needs the MinGW cross compiler and the Windows CUDA runtime on the Linux side;
`.ci/docker/common/install_cuda_windows_cross_compile.sh` installs both.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

major

install_cuda_windows_cross_compile.sh only knows CUDA 12.6, 12.8, 12.9 and 13.0. For cu132 or cu134 it exits with "CUDA version 13.2 is not in the known version map", and the same for 13.4. It is also a script for building the CI image (it runs apt-get as root).

Could the guide describe the setup the workflow uses instead: the MinGW compiler, plus the Windows CUDA runtime for the same CUDA version, with WINDOWS_CUDA_HOME pointing at it?

# The Python side of this wheel carries no CUDA, and install_requirements installs CPU
# torch on Windows, which is what the wheel build and its import checks need. It also
# installs the torchao nightly the wheel declares, which PyPI does not carry.
Invoke-Native { python install_requirements.py }

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

major

I expect the last step of test_cuda_windows.py (the MobileNetV3 run) to fail here once the driver is fixed. install_requirements.py without --example does not install torchvision, and the MobileNetV3 model imports it. The wheel does not depend on it either. I have not seen it fail, because the job stops earlier, at the GPU check.

pre_build_script.sh calls install_requirements.sh --example for the same reason. Could this script do the same?

Comment thread setup.py
f"-DCUDAToolkit_ROOT={windows_cuda_home}",
# Ninja does not generate the Arm Cortex-M Python module's two
# identically named sources as distinct rules.
"-DEXECUTORCH_BUILD_CMSIS_NN_PYBINDS=OFF",

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

minor

This drops the Cortex-M Python module (cmsis_nn) from the CUDA wheel, so Cortex-M lowering cannot work from it. The comment says the module has two identically named sources. I think the real clash is that CMSIS-NN's static library and its test-only shared library both write cmsis-nn.lib. A small test project shows that Ninja error with the two libraries, and setting ARCHIVE_OUTPUT_NAME on the shared one avoids it, so it might make this switch unnecessary. I did not try a full build.

The summary says the export side is identical to the CPU wheel's. Without this module it is not, so it would help to say so in the summary too, not only in the Tests section.

Stacked on #23121, which ships linkable libraries in the Windows wheel.

A Windows user can run a CUDA program but cannot lower one: lowering compiles the model
with a toolchain only the Linux side has, and the CUDA Windows CI job already exports on
Linux and runs on Windows for exactly this reason. What was missing is a wheel that gives
that Windows user the runtime. This adds a Windows CUDA wheel carrying the CUDA delegate
as C++ libraries, the same components the Linux CUDA wheel ships (#21645, #21668), and
nothing CUDA in the Python module, which on Windows has no CUDA program to lower. The
export (AOT) side is unchanged and identical to the CPU wheel's; only the runtime links CUDA.

    lib/executorch_backend_cuda.dll    + .lib   the delegate          executorch::backend_cuda
    lib/executorch_extension_cuda.dll  + .lib   the stream helper     executorch::extension_cuda
    backends/cuda/aoti_cuda_shims.dll            the AOTI shim layer   (runtime DLL of the delegate)
    data/lib/aoti_cuda_shims.lib                 already shipped as the lowering stub; also the
                                                 shim layer's import library here
    extension/pybindings/_C.pyd                  unchanged: no CUDA dependency, no CudaBackend

A C++ application links it exactly as on Linux:

    find_package(executorch REQUIRED COMPONENTS kernels_optimized backend_cuda)
    target_link_libraries(app PRIVATE executorch::runtime executorch::kernels_optimized
                                      executorch::backend_cuda)
    # plus the $<TARGET_RUNTIME_DLLS:app> copy from the CPU PR, which now also brings the
    # stream helper and the shim layer

The wheel's DLLs import no CUDA runtime DLL: CMake links cudart.lib into them on Windows,
and that reaches the driver (nvcuda.dll). The compiled model is different: the library
AOTInductor builds into model.pte imports cudart64_13.dll, which comes from the CUDA
Toolkit's bin directory on PATH, as the guide says. PyPI's nvidia-cuda-runtime has win_amd64
builds that carry that DLL, but a C++ program does not search site-packages, so declaring it
would not make it loadable; the wheel declares no NVIDIA package for Windows.

## How

- Generator. The CUDA toolkit's Visual Studio integration fails compiler identification
  under the ClangCL toolset on some toolkit and Visual Studio pairs (MSB4023 in its targets
  file), and the MSVC toolset cannot compile the Python extension's sources. A Windows CUDA
  build therefore uses Ninja Multi-Config with clang-cl, the compiler the CPU wheel uses,
  and nvcc with cl.exe as its host compiler, when ninja is on PATH. Without it, as in a
  --no-build-isolation source install, the build keeps the Visual Studio generator it always
  used. CMAKE_CUDA_COMPILER is passed only when CUDACXX is unset, so options a user put in
  CUDACXX (such as -allow-unsupported-compiler) still apply. The toolkit is the one
  install_utils already reports, so what compiles is the train packaging declares.
- CUDA 13.2 and newer. Their CCCL stops with an #error under cl.exe's traditional
  preprocessor, so the CUDA sources pass -Xcompiler=/Zc:preprocessor on Windows (13.0
  accepts it too).
- flatc / flatcc byproducts and imported locations share one host executable suffix chosen
  by CMAKE_HOST_WIN32. CMAKE_EXECUTABLE_SUFFIX is the target's: a Zephyr toolchain sets
  .elf, and naming the byproduct after it left Ninja with no rule for the host tool.
- The shims resolve the runtime from executorch.dll rather than a static core. (The PAL's
  strong symbols are in #23121; C++20 for the delegate and shims comes from #23135.)
- The shims keep WINDOWS_EXPORT_ALL_SYMBOLS OFF in a static build, as upstream sets it,
  and turn it ON in the shared layout: the delegate DLL imports the stream guard functions
  (get/set/peek/clearCurrentCUDAStream) from the shim DLL, and those carry no
  AOTI_SHIM_EXPORT, so the annotations alone would not export them.
- #23121 keeps a Windows CUDA build static; this drops that guard, so the pybind preset
  now builds the shared layout with CUDA too, and adds the CUDA import libraries to the
  Windows packaging list.
- extension/cuda links CUDA::cudart before backends/cuda finds the toolkit, so the root
  finds it first; a CPU torch does not define that target.
- The package config adds the stream helper and the shim layer to the delegate's runtime
  DLL set, and reports EXECUTORCH_RUNTIME_DLLS_EXTRA for the pre-3.28 route, since the
  shim layer ships outside lib/.

## CI

- build-wheels-cuda-windows.yml mirrors the Linux CUDA workflow on the Windows wheel
  builder and reuses filter_cuda_matrix.py. pre_build_script.sh drops the blanket Windows
  CUDA=OFF (a CPU row is still forced off by the generic row classification), resolves
  TORCH_CUDA_ARCH_LIST from cuda_arch_list.sh, and writes it to BUILD_ENV_FILE as well as
  GITHUB_ENV: the reusable workflow sources that file after the hook, and its own list had
  replaced the row's (sm_75 in, sm_89 out). It also installs ninja for the generator.
- End-to-end, per published CUDA train, the route a Windows user has, modeled on
  cuda-windows.yml. The trains are read from filter_cuda_matrix.py's SUPPORTED_CUDA_VERSIONS
  (today cu132 and cu134; main stopped publishing cu130 in #23380), the same list the release
  rows are built from, so adding or
  dropping a train there changes what is tested. (install_utils.py's list also has 12.6 for
  source builds, but 12.6 is no longer a published wheel train.) A Linux GPU job exports a
  Windows-target program with that train's torch (`test_cuda_windows.py --export`), and a
  Windows GPU job builds the wheel with that train's nvcc, installs it and runs
  test_cuda_windows.py with the program (test_cuda_windows_e2e.ps1). Trains the images do not
  carry are assembled from NVIDIA's checksummed redistributable archives by
  install_cuda_redist.py, which needs no installer or admin rights and fails on a train it
  has no release for. Each train's Windows job waits only for its own program, so one train
  failing to export does not skip the others. The Windows GPU runners come without an NVIDIA
  driver (nvcuda.dll is missing and cudaGetDeviceCount returns cudaErrorInsufficientDriver),
  so the job first installs the same data-center driver, from the same place, that
  pytorch/pytorch's Windows CUDA smoke tests install (driver_update.bat, 580.88); cu132 and
  cu134 run on it through CUDA's minor-version compatibility, as PyTorch's own wheels do. The
  wheel build's own smoke test has no GPU program and says so.

## Tests

test_cuda_windows.py, the Windows counterpart of test_cuda_linux.py:

- the CUDA DLLs and import libraries ship;
- _C depends on no CUDA DLL and registers no CUDA backend;
- no shipped CUDA DLL imports cudart64_*.dll (only the driver), and no Windows NVIDIA
  requirement is declared;
- every DLL with device code carries exactly the row's GPUs, both directions, with the
  newest also as PTX; the expected list comes from cuda_arch_list.sh keyed by the wheel's
  +cuXYZ tag, not from the TORCH_CUDA_ARCH_LIST the build consumed;
- a C++ application linking executorch::backend_cuda gets the three DLLs copied beside it
  and sees CudaBackend registered;
- a program lowered on Linux for Windows runs through the C++ SDK and matches eager, after
  checking it was exported with the same CUDA train as the wheel (every CUDA 13 train loads
  cudart64_13.dll, so a mismatch would not fail on its own);
- then what a CPU Windows row checks: test_shared_libraries.py (with the three CUDA
  ownership rows), test_cpp_sdk.py and the MobileNetV3/XNNPACK model run. The Arm Cortex-M
  Python module is not built with Ninja, so its install check is left out.

## Test plan

On Windows 11 x64 with an RTX 5080 (sm_120, driver 595.95), VS 2022 BuildTools, CPython 3.12,
and WSL Ubuntu for the export, once per CUDA train. 13.2 and 13.4 are the published trains
and the ones CI covers; 13.0 is no longer published (#23380) but still builds from source, so
it was run too:

| train | Windows nvcc | WSL export torch | wheel | test_cuda_windows.py | Linux-exported program |
|---|---|---|---|---|---|
| 13.0 | 13.0.88 (installer) | 2.14.1+cu130 | executorch-1.6.0+cu130 | 52 checks pass | runs on the GPU, maxdiff 2.4e-7 |
| 13.2 | 13.2.86 (redist) | 2.14.1+cu132 | executorch-1.6.0+cu132 | 52 checks pass | runs on the GPU, maxdiff 2.4e-7 |
| 13.4 | 13.4.59 (redist) | 2.14.0.dev20260810+cu134 | executorch-1.6.0+cu134 | 52 checks pass | runs on the GPU, maxdiff 2.4e-7 |

- Each wheel built with TORCH_CUDA_ARCH_LIST from cuda_arch_list.sh: device code for exactly
  sm_80/86/89/90/100/120, plus sm_120 PTX. 13.4 runs on a driver that reports 13.2, through
  CUDA's minor-version compatibility.
- Before the /Zc:preprocessor change, 13.2 and 13.4 failed to compile the .cu shims with
  CCCL's "traditional preprocessor" #error; 13.0 was unaffected.
- Negative: the cu134 wheel pointed at the 13.0-exported program fails the train check.
- The export used `test_cuda_windows.py --export` from this branch in WSL, with mingw-w64 and
  the Windows CUDA runtime of the same train (13.0 from install_cuda_windows_cross_compile.sh,
  13.2/13.4 from the redist archives).
- The new CI end-to-end jobs could not be run before upload; they follow cuda-windows.yml's
  linux_job_v3 / windows_job pairing and the steps above.
- Linux: the CMake changes are behind WIN32 / MSVC, or pick the same suffix there (empty on a
  Linux host); setup.py's new path runs only on Windows.
@Gasoonjia
Gasoonjia force-pushed the gasoonjia/windows-cuda-sdk-wheel branch from c9578b4 to dd388dd Compare October 3, 2026 23:04

This branch was successfully deployed

1 active deployment
cadence — dd388dd9 Deployed Oct 3, 2026 by Gasoonjia via hifi-op-test / hifi4 #31453
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants