Skip to content

Ship linkable libraries in the Windows CPU wheel - #23121

Open
Gasoonjia wants to merge 1 commit into
mainfrom
gasoonjia/windows-cpu-sdk-wheel
Open

Gasoonjia wants to merge 1 commit into
mainfrom
gasoonjia/windows-cpu-sdk-wheel

Conversation

@Gasoonjia

@Gasoonjia Gasoonjia commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

Ship linkable libraries in the Windows wheel

The Windows wheel shipped one fused Python extension and no linkable libraries, so a C++
application got nothing from it: the headers and the CMake package were there, but every
component a consumer asked for resolved to nothing. Linux and macOS already ship the runtime,
the kernels, the delegates and the thread pool as separate libraries that a C++ application
links directly (#21610, #21771). This does the same for Windows.

Before (Windows x64, CPython 3.12):

extension/pybindings/_C.cp312-win_amd64.pyd   8.27 MB   <- everything fused in here
lib/                                          does not exist

After:

lib/executorch_kernels_optimized.dll          5.02 MB
lib/executorch_backend_xnnpack.dll            2.47 MB
lib/executorch.dll                            0.38 MB
lib/executorch_kernels_quantized.dll          0.22 MB
lib/executorch_threadpool.dll                 0.21 MB
lib/executorch_etdump.dll                     0.05 MB   (loaded by _C; not offered to C++)
lib/*.lib                                               <- import libraries a consumer links
extension/pybindings/_C.cp312-win_amd64.pyd   0.63 MB   <- just the bindings now

How

The three mechanisms that made the shared layout Linux and macOS only now have Windows
equivalents:

what                              ELF / Mach-O                   Windows
exporting the runtime API         default visibility             WINDOWS_EXPORT_ALL_SYMBOLS
keeping a registration-only lib   --no-as-needed / named dylib   /INCLUDE of a per-DLL anchor
finding sibling libraries         $ORIGIN / @loader_path         os.add_dll_directory, and a
                                                                 consumer copies the DLLs
  • Exports. The runtime has no dllexport annotations, which is why CMakeLists.txt refused a
    shared Windows build. Each shipped component DLL now exports all of its symbols, unless
    its target already exports through its own annotations (WINDOWS_EXPORT_ALL_SYMBOLS OFF),
    which it keeps. CMake builds that list from a target's own objects and cannot read an
    archive, so the runtime DLL takes its components as objects rather than through
    /WHOLEARCHIVE. The caller says which archives are part of a DLL's exports (the helper's
    EXPORTED option) instead of the helper inferring it from a property another call sets,
    so the result does not depend on call order. That uses $<COMPILE_ONLY>, new in CMake
    3.27, so a shared Windows build asks for 3.27 with a clear message, the Windows build
    requirement is raised to match, and executorch_shared publishes the cxx_std_20 its
    headers need on Windows. The PAL source is built with its C functions strong there:
    the export list skips weak symbols, so the DLL exported none of the et_pal_* functions
    and clock.h's inline ticks_to_ns() failed to link against it.
  • Retention. A PE import survives only if some symbol from it is referenced, so each
    shipped DLL exports executorch_anchor_<name> and the imported targets carry
    /INCLUDE of it, the counterpart of the Linux --no-as-needed. The name follows the
    shipped file name without any configuration postfix, and is resolved at the end of
    configure and written out literally, so it survives cmake --install (a generator
    expression would be evaluated against the imported target, which has no OUTPUT_NAME)
    and is the same in every configuration of a multi-config build.
  • One registry. A consumer names the runtime's import library ahead of any archive, as on
    Linux, so nothing resolves the registry from a private static copy. Without this the
    optimized kernels registered into their own table: 28 operators visible instead of 242.
  • Loading. A DLL records no search path. The Python entry points that load a DLL depending
    on executorch/lib register that directory first (portable_lib, kernels.quantized,
    codegen.tools, llm.custom_ops.op_tile_crop_aot), computed without resolving symlinks so
    an editable install, where those directories are links, registers its own lib too. A C++
    consumer copies the DLLs beside its executable with $<TARGET_RUNTIME_DLLS>, which the
    C++ guide now shows, including in the quick start.
  • Debug. The DLLs use the C++ library of the configuration the wheel was built in, and a
    consumer in the other configuration would mix the two and corrupt memory. The package
    defines ET_PREBUILT_RELEASE_CRT, or ET_PREBUILT_DEBUG_CRT for a DEBUG=1 build, and the
    runtime headers refuse the mismatched configuration at compile time with a message
    naming the right one, for both single-config (-DCMAKE_BUILD_TYPE) and multi-config
    (--config) generators.
    naming the right one. (A /FAILIFMISMATCH link option would not do: link.exe ignores it
    on the command line and honours it only inside an object file.)
  • cpuinfo. The thread pool DLL exports pthreadpool, so the process has one pool, but keeps
    cpuinfo private. cpuinfo's feature checks (cpuinfo_has_x86_avx2 and the rest) are
    inline reads of cpuinfo_isa, which cpuinfo.h does not declare dllimport, so a DLL can
    only read its own copy; with cpuinfo exported, XNNPACK initialized the thread pool's copy
    and read its own, all zeros, and ran baseline kernels. Each DLL that uses cpuinfo now
    carries all of it. Measured on an x64 machine with AVX2: an XNNPACK-delegated
    2 x Linear(1024) model at batch 256 went from 2.05-2.14 ms to 1.47-1.52 ms per run, with
    the same results.

CUDA is not part of this. A Windows build with CUDA on keeps the static layout it had,
since the Windows CUDA delegate has not been built as a DLL; a follow-up stacked on this
ships it.

Two bugs only the shared layout exposes on Windows are fixed here:

  • The process hung at exit about half the time. Windows terminates worker threads before
    running a DLL's static destructors, so destroying the global thread pool there waited on a
    lock a terminated worker could hold (LdrShutdownProcess -> pthreadpool_destroy -> mtx_lock). The pool is now deliberately not destroyed on Windows; Linux and macOS are
    unchanged.
  • quantized_ops_aot_lib.dll failed to load, which executorch.kernels.quantized swallows,
    so quantized export silently lost its out variants.

The Windows wheel is built without the event tracer (unchanged), so the etdump component
is not offered to C++ there rather than handing a consumer a profiler that records nothing.

The delegates are unchanged too: the Windows wheel ships XNNPACK, as it did before this
change. QNN, OpenVINO and TorchAO stay Linux (and for TorchAO, aarch64) wheel components, and
Core ML and MLX macOS ones. QNN in particular builds on Windows (build-qnn-windows-x64 and
-arm64 pass), but the wheel enables it only where pre_build_script.sh downloads the SDK and
the pybind preset's Linux branch turns it on; bringing it to the Windows wheel means the SDK
download on the Windows builder, its import library and anchor, and tests, which is a
separate change.
The pkg-config file the Linux and macOS wheels ship (#23169) is not generated on Windows
either, since its link flags are GNU-style and a DLL records no search path for them; the
test checks that it is absent.
The pre-3.28 variables route links the import libraries and states C++20, which the Windows
runtime headers need.

Tests

The two wheel suites now run on Windows, as they do on macOS since #21771, wired into
test_windows.py. Their Windows forms ask the platform's own questions:

  • symbols: dumpbin /exports matched against MSVC decorated names, and for code a DLL takes
    from an archive, which the export list cannot name, the import direction: the owner must
    import register_kernels / register_backend from executorch.dll;
  • dependencies: dumpbin /dependents and /imports in place of readelf / otool;
  • loading: LoadLibrary per binary in its own process, which binds every import as ldd -r
    does, from the package and from a relocated copy;
  • paths: every recorded dependency is a bare DLL name;
  • platform tag: the PE machine field of every binary against win_amd64;
  • C++ consumers: built with --config Release, DLLs copied beside the executable, and the
    installed package removed from PATH so nothing resolves it for them;
  • package entry points: the extensions are imported the way a user reaches them (the
    pybindings through portable_lib), with nothing added to the DLL search path, and
    import executorch.kernels.quantized has to register the quantized out variants, so the
    entry points' own registration is what is tested;
  • a consumer in the wrong configuration has to be refused; an application using the
    thread pool has to exit five runs in a row; no DLL may import cpuinfo from another
    (the split state has no numeric symptom, only slower kernels); every program the suites
    build and every Python probe or export they start has a timeout, so a hang fails as a
    named timeout rather than the whole job timing out. Compiler, CMake and binary tool
    invocations are not bounded.

Test plan

On Windows 11 x64, VS 2022 BuildTools with ClangCL, CPython 3.12, from a wheel built with
python setup.py bdist_wheel and installed into a clean environment:

  • .ci/scripts/wheel/test_windows.py: passes, including every check in
    test_shared_libraries.py and test_cpp_sdk.py (single owner of each component, _C contains
    no component, every DLL loads in place and relocated, custom op registers, platform tag;
    version request, 121 of 123 headers compile, documented example builds, runtime alone has
    no kernels, kernels / quantized / XNNPACK consumers match eager PyTorch, delegate fails
    without its component, pre-3.28 route on CMake 3.24, relocation, one shared registry) and
    MobileNetV3 through XNNPACK matching eager.
  • The exit hang: 8 of 20 and 10 of 20 runs of two consumers hung before the thread pool fix,
    0 of 20 after.
  • Linux and macOS: the CMake changes are behind WIN32 / MSVC and the Python changes behind
    sys.platform == "win32"; relying on their wheel rows for confirmation.

@pytorch-bot

pytorch-bot Bot commented Sep 24, 2026 •

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/23121

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure, 1 Pending, 1 Unclassified Failure

As of commit 813c476 with merge base 3794e44 (image):

NEW FAILURE - The following job has failed:

UNCLASSIFIED FAILURE - DrCI could not classify the following job because the workflow did not run on the merge base. The failure may be pre-existing on trunk or introduced by this PR:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-cla meta-cla Bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Sep 24, 2026
@github-actions

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

@Gasoonjia Gasoonjia changed the title Ship linkable libraries in the Windows wheel Ship linkable libraries in the Windows CPU wheel Sep 24, 2026
@Gasoonjia
Gasoonjia force-pushed the gasoonjia/windows-cpu-sdk-wheel branch from 162d4c7 to b10b9e3 Compare September 25, 2026 08:28
Gasoonjia added a commit that referenced this pull request Sep 25, 2026
Stacked on #23121, which ships linkable libraries in the Windows wheel.

A Windows user can run a CUDA program but cannot lower one: lowering compiles the model
with a toolchain only the Linux side has, and the CUDA Windows CI job already exports on
Linux and runs on Windows for exactly this reason. What was missing is a wheel that gives
that Windows user the runtime. This adds a Windows CUDA wheel carrying the CUDA delegate
as C++ libraries, the same components the Linux CUDA wheel ships (#21645, #21668), and
nothing CUDA in the Python module, which on Windows has no CUDA program to lower.

    lib/executorch_backend_cuda.dll    + .lib   the delegate          executorch::backend_cuda
    lib/executorch_extension_cuda.dll  + .lib   the stream helper     executorch::extension_cuda
    backends/cuda/aoti_cuda_shims.dll            the AOTI shim layer   (runtime DLL of the delegate)
    data/lib/aoti_cuda_shims.lib                 already shipped as the lowering stub; also the
                                                 shim layer's import library here
    extension/pybindings/_C.pyd                  unchanged: no CUDA dependency, no CudaBackend

A C++ application links it exactly as on Linux:

    find_package(executorch REQUIRED COMPONENTS kernels_optimized backend_cuda)
    target_link_libraries(app PRIVATE executorch::runtime executorch::kernels_optimized
                                      executorch::backend_cuda)
    # plus the $<TARGET_RUNTIME_DLLS:app> copy from the CPU PR, which now also brings the
    # stream helper and the shim layer

The CUDA runtime is linked into the DLLs (CMake's default on Windows), so the wheel needs
the display driver and not the CUDA toolkit. nvidia publishes no CUDA runtime package for
Windows, so the wheel declares none; the Linux-marked requirements are unchanged.

## How

- Generator. The CUDA toolkit's Visual Studio integration fails compiler identification
  under the ClangCL toolset (MSB4023 in CUDA 13.0.targets), and the MSVC toolset cannot
  compile the Python extension's sources. A Windows CUDA build therefore uses
  Ninja Multi-Config with clang-cl, the compiler the CPU wheel uses, and nvcc with cl.exe
  as its host compiler, the combination the CUDA Windows CI job builds with. The
  multi-config generator keeps the Release/ layout packaging reads. The toolkit is the one
  install_utils already reports, so what compiles is the train packaging declares.
- flatc / flatcc byproducts are declared with the executable suffix. The Visual Studio
  generator ignores them; Ninja on Windows had no rule producing flatc.exe.
- The runtime DLL makes its platform layer strong (ET_PAL_STRONG_SYMBOLS). A DLL resolves
  its own references at link time so a weak PAL cannot be overridden through it anyway,
  and the generated export list skips weak symbols, which left the shims importing
  et_pal_init from nowhere.
- The delegate and shims build as C++20 in the Windows shared layout, which the vendored
  c10 headers need on the MSVC branch; a static build gets that from executorch_core as
  before. The shims resolve the runtime from executorch.dll rather than a static core.
- The shims keep WINDOWS_EXPORT_ALL_SYMBOLS OFF in a static build, as upstream sets it,
  and turn it ON in the shared layout: the delegate DLL imports the stream guard functions
  (get/set/peek/clearCurrentCUDAStream) from the shim DLL, and those carry no
  AOTI_SHIM_EXPORT, so the annotations alone would not export them.
- #23121 keeps a Windows CUDA build static; this drops that guard, so the pybind preset
  now builds the shared layout with CUDA too, and adds the CUDA import libraries to the
  Windows packaging list.
- extension/cuda links CUDA::cudart before backends/cuda finds the toolkit, so the root
  finds it first; a CPU torch does not define that target.
- The package config adds the stream helper and the shim layer to the delegate's runtime
  DLL set, and reports EXECUTORCH_RUNTIME_DLLS_EXTRA for the pre-3.28 route, since the
  shim layer ships outside lib/.
- CI: build-wheels-cuda-windows.yml mirrors the Linux CUDA workflow on the Windows wheel
  builder and reuses filter_cuda_matrix.py. pre_build_script.sh drops the blanket
  Windows CUDA=OFF (a CPU row is still forced off by the generic row classification) and
  resolves TORCH_CUDA_ARCH_LIST from cuda_arch_list.sh, since the Windows builder has no
  env-var script slot. The Arm Cortex-M Python module is off in that build: Ninja does not
  generate its two identically named sources as distinct rules.

## Tests

test_cuda_windows.py, the Windows counterpart of test_cuda_linux.py:

- the CUDA DLLs and import libraries ship;
- _C depends on no CUDA DLL and registers no CUDA backend;
- no CUDA DLL imports cudart64_*.dll (only the driver), and no Windows CUDA requirement
  is declared;
- the shim layer's device code covers the row's TORCH_CUDA_ARCH_LIST;
- a C++ application linking executorch::backend_cuda gets the three DLLs copied beside it
  and sees CudaBackend registered;
- a program lowered on Linux for Windows (`test_cuda_windows.py --export DIR`, which the
  CUDA Windows CI job's export step can produce) runs through the C++ SDK and matches
  eager; without artifacts it prints SKIP, as the Linux aarch64 rows do for execution;
- then everything a CPU Windows row checks: test_shared_libraries.py (now with the three
  CUDA ownership rows) and test_cpp_sdk.py.

## Test plan

On Windows 11 x64 with an RTX 5080 (sm_120), VS 2022 BuildTools, CUDA 13.0, CPython 3.12:

- `CU_VERSION=cu130 TORCH_CUDA_ARCH_LIST=12.0 python setup.py bdist_wheel` ->
  executorch-1.6.0+cu130-cp312-cp312-win_amd64.whl, installed clean.
- In WSL (Ubuntu, torch 2.14.0+cu130, mingw-w64 and the Windows CUDA runtime from
  install_cuda_windows_cross_compile.sh), from this branch:
  `python .ci/scripts/wheel/test_cuda_windows.py --export C:\...\cuda-windows-artifacts`
  -> model.pte (4.9 MB, carrying the Windows kernel DLL) + aoti_cuda_blob.ptd.
- On Windows: `EXECUTORCH_CUDA_WINDOWS_ARTIFACTS=... python test_cuda_windows.py`: every
  check passes, including the Linux-exported program running on the GPU through
  executorch::backend_cuda and matching eager to 2.4e-7.
- CPU regression: the CPU wheel rebuilt from this branch (ClangCL, CUDA off) passes the
  full test_windows.py with the same results as #23121.
- The Windows CUDA CI job (cuda-windows.yml, static llm-release-cuda build) is unaffected:
  every Windows change in backends/cuda is behind EXECUTORCH_BUILD_SHARED.
- Linux: the CMake changes are behind WIN32 / MSVC or evaluate identically there (an
  empty CMAKE_EXECUTABLE_SUFFIX); setup.py's new path runs only on Windows; relying on the
  Linux wheel rows for confirmation.
Comment thread tools/cmake/preset/pybind.cmake
@Gasoonjia
Gasoonjia force-pushed the gasoonjia/windows-cpu-sdk-wheel branch from b10b9e3 to 886edad Compare September 29, 2026 08:23
Gasoonjia added a commit that referenced this pull request Sep 29, 2026
Stacked on #23121, which ships linkable libraries in the Windows wheel.

A Windows user can run a CUDA program but cannot lower one: lowering compiles the model
with a toolchain only the Linux side has, and the CUDA Windows CI job already exports on
Linux and runs on Windows for exactly this reason. What was missing is a wheel that gives
that Windows user the runtime. This adds a Windows CUDA wheel carrying the CUDA delegate
as C++ libraries, the same components the Linux CUDA wheel ships (#21645, #21668), and
nothing CUDA in the Python module, which on Windows has no CUDA program to lower.

    lib/executorch_backend_cuda.dll    + .lib   the delegate          executorch::backend_cuda
    lib/executorch_extension_cuda.dll  + .lib   the stream helper     executorch::extension_cuda
    backends/cuda/aoti_cuda_shims.dll            the AOTI shim layer   (runtime DLL of the delegate)
    data/lib/aoti_cuda_shims.lib                 already shipped as the lowering stub; also the
                                                 shim layer's import library here
    extension/pybindings/_C.pyd                  unchanged: no CUDA dependency, no CudaBackend

A C++ application links it exactly as on Linux:

    find_package(executorch REQUIRED COMPONENTS kernels_optimized backend_cuda)
    target_link_libraries(app PRIVATE executorch::runtime executorch::kernels_optimized
                                      executorch::backend_cuda)
    # plus the $<TARGET_RUNTIME_DLLS:app> copy from the CPU PR, which now also brings the
    # stream helper and the shim layer

The CUDA runtime is linked into the DLLs (CMake's default on Windows), so the wheel needs
the display driver and not the CUDA toolkit. nvidia publishes no CUDA runtime package for
Windows, so the wheel declares none; the Linux-marked requirements are unchanged.

## How

- Generator. The CUDA toolkit's Visual Studio integration fails compiler identification
  under the ClangCL toolset (MSB4023 in CUDA 13.0.targets), and the MSVC toolset cannot
  compile the Python extension's sources. A Windows CUDA build therefore uses
  Ninja Multi-Config with clang-cl, the compiler the CPU wheel uses, and nvcc with cl.exe
  as its host compiler, the combination the CUDA Windows CI job builds with. The
  multi-config generator keeps the Release/ layout packaging reads. The toolkit is the one
  install_utils already reports, so what compiles is the train packaging declares.
- flatc / flatcc byproducts are declared with the executable suffix. The Visual Studio
  generator ignores them; Ninja on Windows had no rule producing flatc.exe.
- The runtime DLL makes its platform layer strong (ET_PAL_STRONG_SYMBOLS). A DLL resolves
  its own references at link time so a weak PAL cannot be overridden through it anyway,
  and the generated export list skips weak symbols, which left the shims importing
  et_pal_init from nowhere.
- The shims resolve the runtime from executorch.dll rather than a static core. (C++20 for
  the delegate and shims on Windows now comes from #23135.)
- The shims keep WINDOWS_EXPORT_ALL_SYMBOLS OFF in a static build, as upstream sets it,
  and turn it ON in the shared layout: the delegate DLL imports the stream guard functions
  (get/set/peek/clearCurrentCUDAStream) from the shim DLL, and those carry no
  AOTI_SHIM_EXPORT, so the annotations alone would not export them.
- #23121 keeps a Windows CUDA build static; this drops that guard, so the pybind preset
  now builds the shared layout with CUDA too, and adds the CUDA import libraries to the
  Windows packaging list.
- extension/cuda links CUDA::cudart before backends/cuda finds the toolkit, so the root
  finds it first; a CPU torch does not define that target.
- The package config adds the stream helper and the shim layer to the delegate's runtime
  DLL set, and reports EXECUTORCH_RUNTIME_DLLS_EXTRA for the pre-3.28 route, since the
  shim layer ships outside lib/.
- CI: build-wheels-cuda-windows.yml mirrors the Linux CUDA workflow on the Windows wheel
  builder and reuses filter_cuda_matrix.py. pre_build_script.sh drops the blanket
  Windows CUDA=OFF (a CPU row is still forced off by the generic row classification) and
  resolves TORCH_CUDA_ARCH_LIST from cuda_arch_list.sh, since the Windows builder has no
  env-var script slot. The Arm Cortex-M Python module is off in that build: Ninja does not
  generate its two identically named sources as distinct rules.

## Tests

test_cuda_windows.py, the Windows counterpart of test_cuda_linux.py:

- the CUDA DLLs and import libraries ship;
- _C depends on no CUDA DLL and registers no CUDA backend;
- no CUDA DLL imports cudart64_*.dll (only the driver), and no Windows CUDA requirement
  is declared;
- the shim layer's device code covers the row's TORCH_CUDA_ARCH_LIST;
- a C++ application linking executorch::backend_cuda gets the three DLLs copied beside it
  and sees CudaBackend registered;
- a program lowered on Linux for Windows (`test_cuda_windows.py --export DIR`, which the
  CUDA Windows CI job's export step can produce) runs through the C++ SDK and matches
  eager; without artifacts it prints SKIP, as the Linux aarch64 rows do for execution;
- then everything a CPU Windows row checks: test_shared_libraries.py (now with the three
  CUDA ownership rows) and test_cpp_sdk.py.

## Test plan

On Windows 11 x64 with an RTX 5080 (sm_120), VS 2022 BuildTools, CUDA 13.0, CPython 3.12:

- `CU_VERSION=cu130 TORCH_CUDA_ARCH_LIST=12.0 python setup.py bdist_wheel` ->
  executorch-1.6.0+cu130-cp312-cp312-win_amd64.whl, installed clean.
- In WSL (Ubuntu, torch 2.14.0+cu130, mingw-w64 and the Windows CUDA runtime from
  install_cuda_windows_cross_compile.sh), from this branch:
  `python .ci/scripts/wheel/test_cuda_windows.py --export C:\...\cuda-windows-artifacts`
  -> model.pte (4.9 MB, carrying the Windows kernel DLL) + aoti_cuda_blob.ptd.
- On Windows: `EXECUTORCH_CUDA_WINDOWS_ARTIFACTS=... python test_cuda_windows.py`: every
  check passes, including the Linux-exported program running on the GPU through
  executorch::backend_cuda and matching eager to 2.4e-7.
- CPU regression: the CPU wheel rebuilt from this branch (ClangCL, CUDA off) passes the
  full test_windows.py with the same results as #23121.
- The Windows CUDA CI job (cuda-windows.yml, static llm-release-cuda build) is unaffected:
  every Windows change in backends/cuda is behind EXECUTORCH_BUILD_SHARED.
- Linux: the CMake changes are behind WIN32 / MSVC or evaluate identically there (an
  empty CMAKE_EXECUTABLE_SUFFIX); setup.py's new path runs only on Windows; relying on the
  Linux wheel rows for confirmation.

@shoumikhin shoumikhin left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Some notes inline.

These are about lines the change does not touch, so they could not be attached to one.

In .ci/scripts/wheel/test_cpp_sdk.py, around line 627:

The description says every program and probe the suites run has a timeout. In the C++ suite only the runs that go through _run_consumer pass one. The runtime-only consumer run, the run without the delegate and the export step call subprocess.run with no timeout, and so do the custom op probe and the parity script in the other suite. If one of them hangs, the job waits for the CI limit instead of failing with a clear timeout. Could those calls use the same timeout, or could the sentence say which runs are bounded?

Comment thread tools/cmake/preset/pybind.cmake
Comment thread tools/cmake/Utils.cmake Outdated
Comment thread runtime/platform/compiler.h
Comment thread tools/cmake/executorch-wheel-config.cmake Outdated
Comment thread tools/cmake/Utils.cmake Outdated
Comment thread CMakeLists.txt
Comment thread tools/cmake/Utils.cmake
Comment thread tools/cmake/Utils.cmake Outdated
Comment thread runtime/platform/compiler.h Outdated
@Gasoonjia
Gasoonjia force-pushed the gasoonjia/windows-cpu-sdk-wheel branch from 24b4561 to bd4757a Compare October 2, 2026 20:00
Gasoonjia added a commit that referenced this pull request Oct 2, 2026
Stacked on #23121, which ships linkable libraries in the Windows wheel.

A Windows user can run a CUDA program but cannot lower one: lowering compiles the model
with a toolchain only the Linux side has, and the CUDA Windows CI job already exports on
Linux and runs on Windows for exactly this reason. What was missing is a wheel that gives
that Windows user the runtime. This adds a Windows CUDA wheel carrying the CUDA delegate
as C++ libraries, the same components the Linux CUDA wheel ships (#21645, #21668), and
nothing CUDA in the Python module, which on Windows has no CUDA program to lower. The
export (AOT) side is unchanged and identical to the CPU wheel's; only the runtime links CUDA.

    lib/executorch_backend_cuda.dll    + .lib   the delegate          executorch::backend_cuda
    lib/executorch_extension_cuda.dll  + .lib   the stream helper     executorch::extension_cuda
    backends/cuda/aoti_cuda_shims.dll            the AOTI shim layer   (runtime DLL of the delegate)
    data/lib/aoti_cuda_shims.lib                 already shipped as the lowering stub; also the
                                                 shim layer's import library here
    extension/pybindings/_C.pyd                  unchanged: no CUDA dependency, no CudaBackend

A C++ application links it exactly as on Linux:

    find_package(executorch REQUIRED COMPONENTS kernels_optimized backend_cuda)
    target_link_libraries(app PRIVATE executorch::runtime executorch::kernels_optimized
                                      executorch::backend_cuda)
    # plus the $<TARGET_RUNTIME_DLLS:app> copy from the CPU PR, which now also brings the
    # stream helper and the shim layer

The wheel's DLLs import no CUDA runtime DLL: CMake links cudart.lib into them on Windows,
and that reaches the driver (nvcuda.dll). The compiled model is different: the library
AOTInductor builds into model.pte imports cudart64_13.dll, which comes from the CUDA
Toolkit's bin directory on PATH, as the guide says. PyPI's nvidia-cuda-runtime has win_amd64
builds that carry that DLL, but a C++ program does not search site-packages, so declaring it
would not make it loadable; the wheel declares no NVIDIA package for Windows.

## How

- Generator. The CUDA toolkit's Visual Studio integration fails compiler identification
  under the ClangCL toolset on some toolkit and Visual Studio pairs (MSB4023 in its targets
  file), and the MSVC toolset cannot compile the Python extension's sources. A Windows CUDA
  build therefore uses Ninja Multi-Config with clang-cl, the compiler the CPU wheel uses,
  and nvcc with cl.exe as its host compiler, when ninja is on PATH. Without it, as in a
  --no-build-isolation source install, the build keeps the Visual Studio generator it always
  used. CMAKE_CUDA_COMPILER is passed only when CUDACXX is unset, so options a user put in
  CUDACXX (such as -allow-unsupported-compiler) still apply. The toolkit is the one
  install_utils already reports, so what compiles is the train packaging declares.
- CUDA 13.2 and newer. Their CCCL stops with an #error under cl.exe's traditional
  preprocessor, so the CUDA sources pass -Xcompiler=/Zc:preprocessor on Windows (13.0
  accepts it too).
- flatc / flatcc byproducts and imported locations share one host executable suffix chosen
  by CMAKE_HOST_WIN32. CMAKE_EXECUTABLE_SUFFIX is the target's: a Zephyr toolchain sets
  .elf, and naming the byproduct after it left Ninja with no rule for the host tool.
- The shims resolve the runtime from executorch.dll rather than a static core. (The PAL's
  strong symbols are in #23121; C++20 for the delegate and shims comes from #23135.)
- The shims keep WINDOWS_EXPORT_ALL_SYMBOLS OFF in a static build, as upstream sets it,
  and turn it ON in the shared layout: the delegate DLL imports the stream guard functions
  (get/set/peek/clearCurrentCUDAStream) from the shim DLL, and those carry no
  AOTI_SHIM_EXPORT, so the annotations alone would not export them.
- #23121 keeps a Windows CUDA build static; this drops that guard, so the pybind preset
  now builds the shared layout with CUDA too, and adds the CUDA import libraries to the
  Windows packaging list.
- extension/cuda links CUDA::cudart before backends/cuda finds the toolkit, so the root
  finds it first; a CPU torch does not define that target.
- The package config adds the stream helper and the shim layer to the delegate's runtime
  DLL set, and reports EXECUTORCH_RUNTIME_DLLS_EXTRA for the pre-3.28 route, since the
  shim layer ships outside lib/.

## CI

- build-wheels-cuda-windows.yml mirrors the Linux CUDA workflow on the Windows wheel
  builder and reuses filter_cuda_matrix.py. pre_build_script.sh drops the blanket Windows
  CUDA=OFF (a CPU row is still forced off by the generic row classification), resolves
  TORCH_CUDA_ARCH_LIST from cuda_arch_list.sh, and writes it to BUILD_ENV_FILE as well as
  GITHUB_ENV: the reusable workflow sources that file after the hook, and its own list had
  replaced the row's (sm_75 in, sm_89 out). It also installs ninja for the generator.
- End-to-end, per published CUDA train, the route a Windows user has, modeled on
  cuda-windows.yml. The trains are read from filter_cuda_matrix.py's SUPPORTED_CUDA_VERSIONS
  (today cu130, cu132, cu134), the same list the release rows are built from, so adding or
  dropping a train there changes what is tested. (install_utils.py's list also has 12.6 for
  source builds, but 12.6 is no longer a published wheel train.) A Linux GPU job exports a
  Windows-target program with that train's torch (`test_cuda_windows.py --export`), and a
  Windows GPU job builds the wheel with that train's nvcc, installs it and runs
  test_cuda_windows.py with the program (test_cuda_windows_e2e.ps1). Trains the images do not
  carry are assembled from NVIDIA's checksummed redistributable archives by
  install_cuda_redist.py, which needs no installer or admin rights and fails on a train it
  has no release for. The wheel build's own smoke test has no GPU program and says so.

## Tests

test_cuda_windows.py, the Windows counterpart of test_cuda_linux.py:

- the CUDA DLLs and import libraries ship;
- _C depends on no CUDA DLL and registers no CUDA backend;
- no shipped CUDA DLL imports cudart64_*.dll (only the driver), and no Windows NVIDIA
  requirement is declared;
- every DLL with device code carries exactly the row's GPUs, both directions, with the
  newest also as PTX; the expected list comes from cuda_arch_list.sh keyed by the wheel's
  +cuXYZ tag, not from the TORCH_CUDA_ARCH_LIST the build consumed;
- a C++ application linking executorch::backend_cuda gets the three DLLs copied beside it
  and sees CudaBackend registered;
- a program lowered on Linux for Windows runs through the C++ SDK and matches eager, after
  checking it was exported with the same CUDA train as the wheel (every CUDA 13 train loads
  cudart64_13.dll, so a mismatch would not fail on its own);
- then what a CPU Windows row checks: test_shared_libraries.py (with the three CUDA
  ownership rows), test_cpp_sdk.py and the MobileNetV3/XNNPACK model run. The Arm Cortex-M
  Python module is not built with Ninja, so its install check is left out.

## Test plan

On Windows 11 x64 with an RTX 5080 (sm_120, driver 595.95), VS 2022 BuildTools, CPython 3.12,
and WSL Ubuntu for the export, once per CUDA train:

| train | Windows nvcc | WSL export torch | wheel | test_cuda_windows.py | Linux-exported program |
|---|---|---|---|---|---|
| 13.0 | 13.0.88 (installer) | 2.14.1+cu130 | executorch-1.6.0+cu130 | 52 checks pass | runs on the GPU, maxdiff 2.4e-7 |
| 13.2 | 13.2.86 (redist) | 2.14.1+cu132 | executorch-1.6.0+cu132 | 52 checks pass | runs on the GPU, maxdiff 2.4e-7 |
| 13.4 | 13.4.59 (redist) | 2.14.0.dev20260810+cu134 | executorch-1.6.0+cu134 | 52 checks pass | runs on the GPU, maxdiff 2.4e-7 |

- Each wheel built with TORCH_CUDA_ARCH_LIST from cuda_arch_list.sh: device code for exactly
  sm_80/86/89/90/100/120, plus sm_120 PTX. 13.4 runs on a driver that reports 13.2, through
  CUDA's minor-version compatibility.
- Before the /Zc:preprocessor change, 13.2 and 13.4 failed to compile the .cu shims with
  CCCL's "traditional preprocessor" #error; 13.0 was unaffected.
- Negative: the cu134 wheel pointed at the 13.0-exported program fails the train check.
- The export used `test_cuda_windows.py --export` from this branch in WSL, with mingw-w64 and
  the Windows CUDA runtime of the same train (13.0 from install_cuda_windows_cross_compile.sh,
  13.2/13.4 from the redist archives).
- The new CI end-to-end jobs could not be run before upload; they follow cuda-windows.yml's
  linux_job_v3 / windows_job pairing and the steps above.
- Linux: the CMake changes are behind WIN32 / MSVC, or pick the same suffix there (empty on a
  Linux host); setup.py's new path runs only on Windows.
Gasoonjia added a commit that referenced this pull request Oct 3, 2026
Stacked on #23121, which ships linkable libraries in the Windows wheel.

A Windows user can run a CUDA program but cannot lower one: lowering compiles the model
with a toolchain only the Linux side has, and the CUDA Windows CI job already exports on
Linux and runs on Windows for exactly this reason. What was missing is a wheel that gives
that Windows user the runtime. This adds a Windows CUDA wheel carrying the CUDA delegate
as C++ libraries, the same components the Linux CUDA wheel ships (#21645, #21668), and
nothing CUDA in the Python module, which on Windows has no CUDA program to lower. The
export (AOT) side is unchanged and identical to the CPU wheel's; only the runtime links CUDA.

    lib/executorch_backend_cuda.dll    + .lib   the delegate          executorch::backend_cuda
    lib/executorch_extension_cuda.dll  + .lib   the stream helper     executorch::extension_cuda
    backends/cuda/aoti_cuda_shims.dll            the AOTI shim layer   (runtime DLL of the delegate)
    data/lib/aoti_cuda_shims.lib                 already shipped as the lowering stub; also the
                                                 shim layer's import library here
    extension/pybindings/_C.pyd                  unchanged: no CUDA dependency, no CudaBackend

A C++ application links it exactly as on Linux:

    find_package(executorch REQUIRED COMPONENTS kernels_optimized backend_cuda)
    target_link_libraries(app PRIVATE executorch::runtime executorch::kernels_optimized
                                      executorch::backend_cuda)
    # plus the $<TARGET_RUNTIME_DLLS:app> copy from the CPU PR, which now also brings the
    # stream helper and the shim layer

The wheel's DLLs import no CUDA runtime DLL: CMake links cudart.lib into them on Windows,
and that reaches the driver (nvcuda.dll). The compiled model is different: the library
AOTInductor builds into model.pte imports cudart64_13.dll, which comes from the CUDA
Toolkit's bin directory on PATH, as the guide says. PyPI's nvidia-cuda-runtime has win_amd64
builds that carry that DLL, but a C++ program does not search site-packages, so declaring it
would not make it loadable; the wheel declares no NVIDIA package for Windows.

## How

- Generator. The CUDA toolkit's Visual Studio integration fails compiler identification
  under the ClangCL toolset on some toolkit and Visual Studio pairs (MSB4023 in its targets
  file), and the MSVC toolset cannot compile the Python extension's sources. A Windows CUDA
  build therefore uses Ninja Multi-Config with clang-cl, the compiler the CPU wheel uses,
  and nvcc with cl.exe as its host compiler, when ninja is on PATH. Without it, as in a
  --no-build-isolation source install, the build keeps the Visual Studio generator it always
  used. CMAKE_CUDA_COMPILER is passed only when CUDACXX is unset, so options a user put in
  CUDACXX (such as -allow-unsupported-compiler) still apply. The toolkit is the one
  install_utils already reports, so what compiles is the train packaging declares.
- CUDA 13.2 and newer. Their CCCL stops with an #error under cl.exe's traditional
  preprocessor, so the CUDA sources pass -Xcompiler=/Zc:preprocessor on Windows (13.0
  accepts it too).
- flatc / flatcc byproducts and imported locations share one host executable suffix chosen
  by CMAKE_HOST_WIN32. CMAKE_EXECUTABLE_SUFFIX is the target's: a Zephyr toolchain sets
  .elf, and naming the byproduct after it left Ninja with no rule for the host tool.
- The shims resolve the runtime from executorch.dll rather than a static core. (The PAL's
  strong symbols are in #23121; C++20 for the delegate and shims comes from #23135.)
- The shims keep WINDOWS_EXPORT_ALL_SYMBOLS OFF in a static build, as upstream sets it,
  and turn it ON in the shared layout: the delegate DLL imports the stream guard functions
  (get/set/peek/clearCurrentCUDAStream) from the shim DLL, and those carry no
  AOTI_SHIM_EXPORT, so the annotations alone would not export them.
- #23121 keeps a Windows CUDA build static; this drops that guard, so the pybind preset
  now builds the shared layout with CUDA too, and adds the CUDA import libraries to the
  Windows packaging list.
- extension/cuda links CUDA::cudart before backends/cuda finds the toolkit, so the root
  finds it first; a CPU torch does not define that target.
- The package config adds the stream helper and the shim layer to the delegate's runtime
  DLL set, and reports EXECUTORCH_RUNTIME_DLLS_EXTRA for the pre-3.28 route, since the
  shim layer ships outside lib/.

## CI

- build-wheels-cuda-windows.yml mirrors the Linux CUDA workflow on the Windows wheel
  builder and reuses filter_cuda_matrix.py. pre_build_script.sh drops the blanket Windows
  CUDA=OFF (a CPU row is still forced off by the generic row classification), resolves
  TORCH_CUDA_ARCH_LIST from cuda_arch_list.sh, and writes it to BUILD_ENV_FILE as well as
  GITHUB_ENV: the reusable workflow sources that file after the hook, and its own list had
  replaced the row's (sm_75 in, sm_89 out). It also installs ninja for the generator.
- End-to-end, per published CUDA train, the route a Windows user has, modeled on
  cuda-windows.yml. The trains are read from filter_cuda_matrix.py's SUPPORTED_CUDA_VERSIONS
  (today cu130, cu132, cu134), the same list the release rows are built from, so adding or
  dropping a train there changes what is tested. (install_utils.py's list also has 12.6 for
  source builds, but 12.6 is no longer a published wheel train.) A Linux GPU job exports a
  Windows-target program with that train's torch (`test_cuda_windows.py --export`), and a
  Windows GPU job builds the wheel with that train's nvcc, installs it and runs
  test_cuda_windows.py with the program (test_cuda_windows_e2e.ps1). Trains the images do not
  carry are assembled from NVIDIA's checksummed redistributable archives by
  install_cuda_redist.py, which needs no installer or admin rights and fails on a train it
  has no release for. Each train's Windows job waits only for its own program, so one train
  failing to export does not skip the others. The wheel build's own smoke test has no GPU
  program and says so.

## Tests

test_cuda_windows.py, the Windows counterpart of test_cuda_linux.py:

- the CUDA DLLs and import libraries ship;
- _C depends on no CUDA DLL and registers no CUDA backend;
- no shipped CUDA DLL imports cudart64_*.dll (only the driver), and no Windows NVIDIA
  requirement is declared;
- every DLL with device code carries exactly the row's GPUs, both directions, with the
  newest also as PTX; the expected list comes from cuda_arch_list.sh keyed by the wheel's
  +cuXYZ tag, not from the TORCH_CUDA_ARCH_LIST the build consumed;
- a C++ application linking executorch::backend_cuda gets the three DLLs copied beside it
  and sees CudaBackend registered;
- a program lowered on Linux for Windows runs through the C++ SDK and matches eager, after
  checking it was exported with the same CUDA train as the wheel (every CUDA 13 train loads
  cudart64_13.dll, so a mismatch would not fail on its own);
- then what a CPU Windows row checks: test_shared_libraries.py (with the three CUDA
  ownership rows), test_cpp_sdk.py and the MobileNetV3/XNNPACK model run. The Arm Cortex-M
  Python module is not built with Ninja, so its install check is left out.

## Test plan

On Windows 11 x64 with an RTX 5080 (sm_120, driver 595.95), VS 2022 BuildTools, CPython 3.12,
and WSL Ubuntu for the export, once per CUDA train:

| train | Windows nvcc | WSL export torch | wheel | test_cuda_windows.py | Linux-exported program |
|---|---|---|---|---|---|
| 13.0 | 13.0.88 (installer) | 2.14.1+cu130 | executorch-1.6.0+cu130 | 52 checks pass | runs on the GPU, maxdiff 2.4e-7 |
| 13.2 | 13.2.86 (redist) | 2.14.1+cu132 | executorch-1.6.0+cu132 | 52 checks pass | runs on the GPU, maxdiff 2.4e-7 |
| 13.4 | 13.4.59 (redist) | 2.14.0.dev20260810+cu134 | executorch-1.6.0+cu134 | 52 checks pass | runs on the GPU, maxdiff 2.4e-7 |

- Each wheel built with TORCH_CUDA_ARCH_LIST from cuda_arch_list.sh: device code for exactly
  sm_80/86/89/90/100/120, plus sm_120 PTX. 13.4 runs on a driver that reports 13.2, through
  CUDA's minor-version compatibility.
- Before the /Zc:preprocessor change, 13.2 and 13.4 failed to compile the .cu shims with
  CCCL's "traditional preprocessor" #error; 13.0 was unaffected.
- Negative: the cu134 wheel pointed at the 13.0-exported program fails the train check.
- The export used `test_cuda_windows.py --export` from this branch in WSL, with mingw-w64 and
  the Windows CUDA runtime of the same train (13.0 from install_cuda_windows_cross_compile.sh,
  13.2/13.4 from the redist archives).
- The new CI end-to-end jobs could not be run before upload; they follow cuda-windows.yml's
  linux_job_v3 / windows_job pairing and the steps above.
- Linux: the CMake changes are behind WIN32 / MSVC, or pick the same suffix there (empty on a
  Linux host); setup.py's new path runs only on Windows.
The Windows wheel shipped one fused Python extension and no linkable libraries, so a C++
application got nothing from it: the headers and the CMake package were there, but every
component a consumer asked for resolved to nothing. Linux and macOS already ship the runtime,
the kernels, the delegates and the thread pool as separate libraries that a C++ application
links directly (#21610, #21771). This does the same for Windows.

Before (Windows x64, CPython 3.12):

    extension/pybindings/_C.cp312-win_amd64.pyd   8.27 MB   <- everything fused in here
    lib/                                          does not exist

After:

    lib/executorch_kernels_optimized.dll          5.02 MB
    lib/executorch_backend_xnnpack.dll            2.47 MB
    lib/executorch.dll                            0.38 MB
    lib/executorch_kernels_quantized.dll          0.22 MB
    lib/executorch_threadpool.dll                 0.21 MB
    lib/executorch_etdump.dll                     0.05 MB   (loaded by _C; not offered to C++)
    lib/*.lib                                               <- import libraries a consumer links
    extension/pybindings/_C.cp312-win_amd64.pyd   0.63 MB   <- just the bindings now

## How

The three mechanisms that made the shared layout Linux and macOS only now have Windows
equivalents:

    what                              ELF / Mach-O                   Windows
    exporting the runtime API         default visibility             WINDOWS_EXPORT_ALL_SYMBOLS
    keeping a registration-only lib   --no-as-needed / named dylib   /INCLUDE of a per-DLL anchor
    finding sibling libraries         $ORIGIN / @loader_path         os.add_dll_directory, and a
                                                                     consumer copies the DLLs

- Exports. The runtime has no dllexport annotations, which is why CMakeLists.txt refused a
  shared Windows build. Each shipped component DLL now exports all of its symbols, unless
  its target already exports through its own annotations (WINDOWS_EXPORT_ALL_SYMBOLS OFF),
  which it keeps. CMake builds that list from a target's own objects and cannot read an
  archive, so the runtime DLL takes its components as objects rather than through
  /WHOLEARCHIVE. The caller says which archives are part of a DLL's exports (the helper's
  EXPORTED option) instead of the helper inferring it from a property another call sets,
  so the result does not depend on call order. That uses $<COMPILE_ONLY>, new in CMake
  3.27, so a shared Windows build asks for 3.27 with a clear message, the Windows build
  requirement is raised to match, and executorch_shared publishes the cxx_std_20 its
  headers need on Windows. The PAL source is built with its C functions strong there:
  the export list skips weak symbols, so the DLL exported none of the et_pal_* functions
  and clock.h's inline ticks_to_ns() failed to link against it.
- Retention. A PE import survives only if some symbol from it is referenced, so each
  shipped DLL exports `executorch_anchor_<name>` and the imported targets carry
  `/INCLUDE` of it, the counterpart of the Linux `--no-as-needed`. The name follows the
  shipped file name without any configuration postfix, and is resolved at the end of
  configure and written out literally, so it survives `cmake --install` (a generator
  expression would be evaluated against the imported target, which has no OUTPUT_NAME)
  and is the same in every configuration of a multi-config build.
- One registry. A consumer names the runtime's import library ahead of any archive, as on
  Linux, so nothing resolves the registry from a private static copy. Without this the
  optimized kernels registered into their own table: 28 operators visible instead of 242.
- Loading. A DLL records no search path. The Python entry points that load a DLL depending
  on executorch/lib register that directory first (portable_lib, kernels.quantized,
  codegen.tools, llm.custom_ops.op_tile_crop_aot), computed without resolving symlinks so
  an editable install, where those directories are links, registers its own lib too. A C++
  consumer copies the DLLs beside its executable with `$<TARGET_RUNTIME_DLLS>`, which the
  C++ guide now shows, including in the quick start.
- Debug. The DLLs use the C++ library of the configuration the wheel was built in, and a
  consumer in the other configuration would mix the two and corrupt memory. The package
  defines ET_PREBUILT_RELEASE_CRT, or ET_PREBUILT_DEBUG_CRT for a DEBUG=1 build, and the
  runtime headers refuse the mismatched configuration at compile time with a message
  naming the right one, for both single-config (-DCMAKE_BUILD_TYPE) and multi-config
  (--config) generators.
  naming the right one. (A /FAILIFMISMATCH link option would not do: link.exe ignores it
  on the command line and honours it only inside an object file.)
- cpuinfo. The thread pool DLL exports pthreadpool, so the process has one pool, but keeps
  cpuinfo private. cpuinfo's feature checks (cpuinfo_has_x86_avx2 and the rest) are
  inline reads of cpuinfo_isa, which cpuinfo.h does not declare dllimport, so a DLL can
  only read its own copy; with cpuinfo exported, XNNPACK initialized the thread pool's copy
  and read its own, all zeros, and ran baseline kernels. Each DLL that uses cpuinfo now
  carries all of it. Measured on an x64 machine with AVX2: an XNNPACK-delegated
  2 x Linear(1024) model at batch 256 went from 2.05-2.14 ms to 1.47-1.52 ms per run, with
  the same results.

CUDA is not part of this. A Windows build with CUDA on keeps the static layout it had,
since the Windows CUDA delegate has not been built as a DLL; a follow-up stacked on this
ships it.

Two bugs only the shared layout exposes on Windows are fixed here:

- The process hung at exit about half the time. Windows terminates worker threads before
  running a DLL's static destructors, so destroying the global thread pool there waited on a
  lock a terminated worker could hold (`LdrShutdownProcess -> pthreadpool_destroy ->
  mtx_lock`). The pool is now deliberately not destroyed on Windows; Linux and macOS are
  unchanged.
- `quantized_ops_aot_lib.dll` failed to load, which `executorch.kernels.quantized` swallows,
  so quantized export silently lost its out variants.

The Windows wheel is built without the event tracer (unchanged), so the `etdump` component
is not offered to C++ there rather than handing a consumer a profiler that records nothing.

The delegates are unchanged too: the Windows wheel ships XNNPACK, as it did before this
change. QNN, OpenVINO and TorchAO stay Linux (and for TorchAO, aarch64) wheel components, and
Core ML and MLX macOS ones. QNN in particular builds on Windows (build-qnn-windows-x64 and
-arm64 pass), but the wheel enables it only where pre_build_script.sh downloads the SDK and
the pybind preset's Linux branch turns it on; bringing it to the Windows wheel means the SDK
download on the Windows builder, its import library and anchor, and tests, which is a
separate change.
The pkg-config file the Linux and macOS wheels ship (#23169) is not generated on Windows
either, since its link flags are GNU-style and a DLL records no search path for them; the
test checks that it is absent.
The pre-3.28 variables route links the import libraries and states C++20, which the Windows
runtime headers need.

## Tests

The two wheel suites now run on Windows, as they do on macOS since #21771, wired into
test_windows.py. Their Windows forms ask the platform's own questions:

- symbols: dumpbin /exports matched against MSVC decorated names, and for code a DLL takes
  from an archive, which the export list cannot name, the import direction: the owner must
  import `register_kernels` / `register_backend` from executorch.dll;
- dependencies: dumpbin /dependents and /imports in place of readelf / otool;
- loading: LoadLibrary per binary in its own process, which binds every import as ldd -r
  does, from the package and from a relocated copy;
- paths: every recorded dependency is a bare DLL name;
- platform tag: the PE machine field of every binary against win_amd64;
- C++ consumers: built with --config Release, DLLs copied beside the executable, and the
  installed package removed from PATH so nothing resolves it for them;
- package entry points: the extensions are imported the way a user reaches them (the
  pybindings through portable_lib), with nothing added to the DLL search path, and
  `import executorch.kernels.quantized` has to register the quantized out variants, so the
  entry points' own registration is what is tested;
- a consumer in the wrong configuration has to be refused; an application using the
  thread pool has to exit five runs in a row; no DLL may import cpuinfo from another
  (the split state has no numeric symptom, only slower kernels); every program the suites
  build and every Python probe or export they start has a timeout, so a hang fails as a
  named timeout rather than the whole job timing out. Compiler, CMake and binary tool
  invocations are not bounded.

## Test plan

On Windows 11 x64, VS 2022 BuildTools with ClangCL, CPython 3.12, from a wheel built with
`python setup.py bdist_wheel` and installed into a clean environment:

- `.ci/scripts/wheel/test_windows.py`: passes, including every check in
  test_shared_libraries.py and test_cpp_sdk.py (single owner of each component, _C contains
  no component, every DLL loads in place and relocated, custom op registers, platform tag;
  version request, 121 of 123 headers compile, documented example builds, runtime alone has
  no kernels, kernels / quantized / XNNPACK consumers match eager PyTorch, delegate fails
  without its component, pre-3.28 route on CMake 3.24, relocation, one shared registry) and
  MobileNetV3 through XNNPACK matching eager.
- The exit hang: 8 of 20 and 10 of 20 runs of two consumers hung before the thread pool fix,
  0 of 20 after.
- Linux and macOS: the CMake changes are behind WIN32 / MSVC and the Python changes behind
  sys.platform == "win32"; relying on their wheel rows for confirmation.
@Gasoonjia
Gasoonjia force-pushed the gasoonjia/windows-cpu-sdk-wheel branch from bd4757a to 813c476 Compare October 3, 2026 02:23
Gasoonjia added a commit that referenced this pull request Oct 3, 2026
Stacked on #23121, which ships linkable libraries in the Windows wheel.

A Windows user can run a CUDA program but cannot lower one: lowering compiles the model
with a toolchain only the Linux side has, and the CUDA Windows CI job already exports on
Linux and runs on Windows for exactly this reason. What was missing is a wheel that gives
that Windows user the runtime. This adds a Windows CUDA wheel carrying the CUDA delegate
as C++ libraries, the same components the Linux CUDA wheel ships (#21645, #21668), and
nothing CUDA in the Python module, which on Windows has no CUDA program to lower. The
export (AOT) side is unchanged and identical to the CPU wheel's; only the runtime links CUDA.

    lib/executorch_backend_cuda.dll    + .lib   the delegate          executorch::backend_cuda
    lib/executorch_extension_cuda.dll  + .lib   the stream helper     executorch::extension_cuda
    backends/cuda/aoti_cuda_shims.dll            the AOTI shim layer   (runtime DLL of the delegate)
    data/lib/aoti_cuda_shims.lib                 already shipped as the lowering stub; also the
                                                 shim layer's import library here
    extension/pybindings/_C.pyd                  unchanged: no CUDA dependency, no CudaBackend

A C++ application links it exactly as on Linux:

    find_package(executorch REQUIRED COMPONENTS kernels_optimized backend_cuda)
    target_link_libraries(app PRIVATE executorch::runtime executorch::kernels_optimized
                                      executorch::backend_cuda)
    # plus the $<TARGET_RUNTIME_DLLS:app> copy from the CPU PR, which now also brings the
    # stream helper and the shim layer

The wheel's DLLs import no CUDA runtime DLL: CMake links cudart.lib into them on Windows,
and that reaches the driver (nvcuda.dll). The compiled model is different: the library
AOTInductor builds into model.pte imports cudart64_13.dll, which comes from the CUDA
Toolkit's bin directory on PATH, as the guide says. PyPI's nvidia-cuda-runtime has win_amd64
builds that carry that DLL, but a C++ program does not search site-packages, so declaring it
would not make it loadable; the wheel declares no NVIDIA package for Windows.

## How

- Generator. The CUDA toolkit's Visual Studio integration fails compiler identification
  under the ClangCL toolset on some toolkit and Visual Studio pairs (MSB4023 in its targets
  file), and the MSVC toolset cannot compile the Python extension's sources. A Windows CUDA
  build therefore uses Ninja Multi-Config with clang-cl, the compiler the CPU wheel uses,
  and nvcc with cl.exe as its host compiler, when ninja is on PATH. Without it, as in a
  --no-build-isolation source install, the build keeps the Visual Studio generator it always
  used. CMAKE_CUDA_COMPILER is passed only when CUDACXX is unset, so options a user put in
  CUDACXX (such as -allow-unsupported-compiler) still apply. The toolkit is the one
  install_utils already reports, so what compiles is the train packaging declares.
- CUDA 13.2 and newer. Their CCCL stops with an #error under cl.exe's traditional
  preprocessor, so the CUDA sources pass -Xcompiler=/Zc:preprocessor on Windows (13.0
  accepts it too).
- flatc / flatcc byproducts and imported locations share one host executable suffix chosen
  by CMAKE_HOST_WIN32. CMAKE_EXECUTABLE_SUFFIX is the target's: a Zephyr toolchain sets
  .elf, and naming the byproduct after it left Ninja with no rule for the host tool.
- The shims resolve the runtime from executorch.dll rather than a static core. (The PAL's
  strong symbols are in #23121; C++20 for the delegate and shims comes from #23135.)
- The shims keep WINDOWS_EXPORT_ALL_SYMBOLS OFF in a static build, as upstream sets it,
  and turn it ON in the shared layout: the delegate DLL imports the stream guard functions
  (get/set/peek/clearCurrentCUDAStream) from the shim DLL, and those carry no
  AOTI_SHIM_EXPORT, so the annotations alone would not export them.
- #23121 keeps a Windows CUDA build static; this drops that guard, so the pybind preset
  now builds the shared layout with CUDA too, and adds the CUDA import libraries to the
  Windows packaging list.
- extension/cuda links CUDA::cudart before backends/cuda finds the toolkit, so the root
  finds it first; a CPU torch does not define that target.
- The package config adds the stream helper and the shim layer to the delegate's runtime
  DLL set, and reports EXECUTORCH_RUNTIME_DLLS_EXTRA for the pre-3.28 route, since the
  shim layer ships outside lib/.

## CI

- build-wheels-cuda-windows.yml mirrors the Linux CUDA workflow on the Windows wheel
  builder and reuses filter_cuda_matrix.py. pre_build_script.sh drops the blanket Windows
  CUDA=OFF (a CPU row is still forced off by the generic row classification), resolves
  TORCH_CUDA_ARCH_LIST from cuda_arch_list.sh, and writes it to BUILD_ENV_FILE as well as
  GITHUB_ENV: the reusable workflow sources that file after the hook, and its own list had
  replaced the row's (sm_75 in, sm_89 out). It also installs ninja for the generator.
- End-to-end, per published CUDA train, the route a Windows user has, modeled on
  cuda-windows.yml. The trains are read from filter_cuda_matrix.py's SUPPORTED_CUDA_VERSIONS
  (today cu132 and cu134; main stopped publishing cu130 in #23380), the same list the release
  rows are built from, so adding or
  dropping a train there changes what is tested. (install_utils.py's list also has 12.6 for
  source builds, but 12.6 is no longer a published wheel train.) A Linux GPU job exports a
  Windows-target program with that train's torch (`test_cuda_windows.py --export`), and a
  Windows GPU job builds the wheel with that train's nvcc, installs it and runs
  test_cuda_windows.py with the program (test_cuda_windows_e2e.ps1). Trains the images do not
  carry are assembled from NVIDIA's checksummed redistributable archives by
  install_cuda_redist.py, which needs no installer or admin rights and fails on a train it
  has no release for. Each train's Windows job waits only for its own program, so one train
  failing to export does not skip the others. The wheel build's own smoke test has no GPU
  program and says so.

## Tests

test_cuda_windows.py, the Windows counterpart of test_cuda_linux.py:

- the CUDA DLLs and import libraries ship;
- _C depends on no CUDA DLL and registers no CUDA backend;
- no shipped CUDA DLL imports cudart64_*.dll (only the driver), and no Windows NVIDIA
  requirement is declared;
- every DLL with device code carries exactly the row's GPUs, both directions, with the
  newest also as PTX; the expected list comes from cuda_arch_list.sh keyed by the wheel's
  +cuXYZ tag, not from the TORCH_CUDA_ARCH_LIST the build consumed;
- a C++ application linking executorch::backend_cuda gets the three DLLs copied beside it
  and sees CudaBackend registered;
- a program lowered on Linux for Windows runs through the C++ SDK and matches eager, after
  checking it was exported with the same CUDA train as the wheel (every CUDA 13 train loads
  cudart64_13.dll, so a mismatch would not fail on its own);
- then what a CPU Windows row checks: test_shared_libraries.py (with the three CUDA
  ownership rows), test_cpp_sdk.py and the MobileNetV3/XNNPACK model run. The Arm Cortex-M
  Python module is not built with Ninja, so its install check is left out.

## Test plan

On Windows 11 x64 with an RTX 5080 (sm_120, driver 595.95), VS 2022 BuildTools, CPython 3.12,
and WSL Ubuntu for the export, once per CUDA train. 13.2 and 13.4 are the published trains
and the ones CI covers; 13.0 is no longer published (#23380) but still builds from source, so
it was run too:

| train | Windows nvcc | WSL export torch | wheel | test_cuda_windows.py | Linux-exported program |
|---|---|---|---|---|---|
| 13.0 | 13.0.88 (installer) | 2.14.1+cu130 | executorch-1.6.0+cu130 | 52 checks pass | runs on the GPU, maxdiff 2.4e-7 |
| 13.2 | 13.2.86 (redist) | 2.14.1+cu132 | executorch-1.6.0+cu132 | 52 checks pass | runs on the GPU, maxdiff 2.4e-7 |
| 13.4 | 13.4.59 (redist) | 2.14.0.dev20260810+cu134 | executorch-1.6.0+cu134 | 52 checks pass | runs on the GPU, maxdiff 2.4e-7 |

- Each wheel built with TORCH_CUDA_ARCH_LIST from cuda_arch_list.sh: device code for exactly
  sm_80/86/89/90/100/120, plus sm_120 PTX. 13.4 runs on a driver that reports 13.2, through
  CUDA's minor-version compatibility.
- Before the /Zc:preprocessor change, 13.2 and 13.4 failed to compile the .cu shims with
  CCCL's "traditional preprocessor" #error; 13.0 was unaffected.
- Negative: the cu134 wheel pointed at the 13.0-exported program fails the train check.
- The export used `test_cuda_windows.py --export` from this branch in WSL, with mingw-w64 and
  the Windows CUDA runtime of the same train (13.0 from install_cuda_windows_cross_compile.sh,
  13.2/13.4 from the redist archives).
- The new CI end-to-end jobs could not be run before upload; they follow cuda-windows.yml's
  linux_job_v3 / windows_job pairing and the steps above.
- Linux: the CMake changes are behind WIN32 / MSVC, or pick the same suffix there (empty on a
  Linux host); setup.py's new path runs only on Windows.
Gasoonjia added a commit that referenced this pull request Oct 3, 2026
Stacked on #23121, which ships linkable libraries in the Windows wheel.

A Windows user can run a CUDA program but cannot lower one: lowering compiles the model
with a toolchain only the Linux side has, and the CUDA Windows CI job already exports on
Linux and runs on Windows for exactly this reason. What was missing is a wheel that gives
that Windows user the runtime. This adds a Windows CUDA wheel carrying the CUDA delegate
as C++ libraries, the same components the Linux CUDA wheel ships (#21645, #21668), and
nothing CUDA in the Python module, which on Windows has no CUDA program to lower. The
export (AOT) side is unchanged and identical to the CPU wheel's; only the runtime links CUDA.

    lib/executorch_backend_cuda.dll    + .lib   the delegate          executorch::backend_cuda
    lib/executorch_extension_cuda.dll  + .lib   the stream helper     executorch::extension_cuda
    backends/cuda/aoti_cuda_shims.dll            the AOTI shim layer   (runtime DLL of the delegate)
    data/lib/aoti_cuda_shims.lib                 already shipped as the lowering stub; also the
                                                 shim layer's import library here
    extension/pybindings/_C.pyd                  unchanged: no CUDA dependency, no CudaBackend

A C++ application links it exactly as on Linux:

    find_package(executorch REQUIRED COMPONENTS kernels_optimized backend_cuda)
    target_link_libraries(app PRIVATE executorch::runtime executorch::kernels_optimized
                                      executorch::backend_cuda)
    # plus the $<TARGET_RUNTIME_DLLS:app> copy from the CPU PR, which now also brings the
    # stream helper and the shim layer

The wheel's DLLs import no CUDA runtime DLL: CMake links cudart.lib into them on Windows,
and that reaches the driver (nvcuda.dll). The compiled model is different: the library
AOTInductor builds into model.pte imports cudart64_13.dll, which comes from the CUDA
Toolkit's bin directory on PATH, as the guide says. PyPI's nvidia-cuda-runtime has win_amd64
builds that carry that DLL, but a C++ program does not search site-packages, so declaring it
would not make it loadable; the wheel declares no NVIDIA package for Windows.

## How

- Generator. The CUDA toolkit's Visual Studio integration fails compiler identification
  under the ClangCL toolset on some toolkit and Visual Studio pairs (MSB4023 in its targets
  file), and the MSVC toolset cannot compile the Python extension's sources. A Windows CUDA
  build therefore uses Ninja Multi-Config with clang-cl, the compiler the CPU wheel uses,
  and nvcc with cl.exe as its host compiler, when ninja is on PATH. Without it, as in a
  --no-build-isolation source install, the build keeps the Visual Studio generator it always
  used. CMAKE_CUDA_COMPILER is passed only when CUDACXX is unset, so options a user put in
  CUDACXX (such as -allow-unsupported-compiler) still apply. The toolkit is the one
  install_utils already reports, so what compiles is the train packaging declares.
- CUDA 13.2 and newer. Their CCCL stops with an #error under cl.exe's traditional
  preprocessor, so the CUDA sources pass -Xcompiler=/Zc:preprocessor on Windows (13.0
  accepts it too).
- flatc / flatcc byproducts and imported locations share one host executable suffix chosen
  by CMAKE_HOST_WIN32. CMAKE_EXECUTABLE_SUFFIX is the target's: a Zephyr toolchain sets
  .elf, and naming the byproduct after it left Ninja with no rule for the host tool.
- The shims resolve the runtime from executorch.dll rather than a static core. (The PAL's
  strong symbols are in #23121; C++20 for the delegate and shims comes from #23135.)
- The shims keep WINDOWS_EXPORT_ALL_SYMBOLS OFF in a static build, as upstream sets it,
  and turn it ON in the shared layout: the delegate DLL imports the stream guard functions
  (get/set/peek/clearCurrentCUDAStream) from the shim DLL, and those carry no
  AOTI_SHIM_EXPORT, so the annotations alone would not export them.
- #23121 keeps a Windows CUDA build static; this drops that guard, so the pybind preset
  now builds the shared layout with CUDA too, and adds the CUDA import libraries to the
  Windows packaging list.
- extension/cuda links CUDA::cudart before backends/cuda finds the toolkit, so the root
  finds it first; a CPU torch does not define that target.
- The package config adds the stream helper and the shim layer to the delegate's runtime
  DLL set, and reports EXECUTORCH_RUNTIME_DLLS_EXTRA for the pre-3.28 route, since the
  shim layer ships outside lib/.

## CI

- build-wheels-cuda-windows.yml mirrors the Linux CUDA workflow on the Windows wheel
  builder and reuses filter_cuda_matrix.py. pre_build_script.sh drops the blanket Windows
  CUDA=OFF (a CPU row is still forced off by the generic row classification), resolves
  TORCH_CUDA_ARCH_LIST from cuda_arch_list.sh, and writes it to BUILD_ENV_FILE as well as
  GITHUB_ENV: the reusable workflow sources that file after the hook, and its own list had
  replaced the row's (sm_75 in, sm_89 out). It also installs ninja for the generator.
- End-to-end, per published CUDA train, the route a Windows user has, modeled on
  cuda-windows.yml. The trains are read from filter_cuda_matrix.py's SUPPORTED_CUDA_VERSIONS
  (today cu132 and cu134; main stopped publishing cu130 in #23380), the same list the release
  rows are built from, so adding or
  dropping a train there changes what is tested. (install_utils.py's list also has 12.6 for
  source builds, but 12.6 is no longer a published wheel train.) A Linux GPU job exports a
  Windows-target program with that train's torch (`test_cuda_windows.py --export`), and a
  Windows GPU job builds the wheel with that train's nvcc, installs it and runs
  test_cuda_windows.py with the program (test_cuda_windows_e2e.ps1). Trains the images do not
  carry are assembled from NVIDIA's checksummed redistributable archives by
  install_cuda_redist.py, which needs no installer or admin rights and fails on a train it
  has no release for. Each train's Windows job waits only for its own program, so one train
  failing to export does not skip the others. The wheel build's own smoke test has no GPU
  program and says so.

## Tests

test_cuda_windows.py, the Windows counterpart of test_cuda_linux.py:

- the CUDA DLLs and import libraries ship;
- _C depends on no CUDA DLL and registers no CUDA backend;
- no shipped CUDA DLL imports cudart64_*.dll (only the driver), and no Windows NVIDIA
  requirement is declared;
- every DLL with device code carries exactly the row's GPUs, both directions, with the
  newest also as PTX; the expected list comes from cuda_arch_list.sh keyed by the wheel's
  +cuXYZ tag, not from the TORCH_CUDA_ARCH_LIST the build consumed;
- a C++ application linking executorch::backend_cuda gets the three DLLs copied beside it
  and sees CudaBackend registered;
- a program lowered on Linux for Windows runs through the C++ SDK and matches eager, after
  checking it was exported with the same CUDA train as the wheel (every CUDA 13 train loads
  cudart64_13.dll, so a mismatch would not fail on its own);
- then what a CPU Windows row checks: test_shared_libraries.py (with the three CUDA
  ownership rows), test_cpp_sdk.py and the MobileNetV3/XNNPACK model run. The Arm Cortex-M
  Python module is not built with Ninja, so its install check is left out.

## Test plan

On Windows 11 x64 with an RTX 5080 (sm_120, driver 595.95), VS 2022 BuildTools, CPython 3.12,
and WSL Ubuntu for the export, once per CUDA train. 13.2 and 13.4 are the published trains
and the ones CI covers; 13.0 is no longer published (#23380) but still builds from source, so
it was run too:

| train | Windows nvcc | WSL export torch | wheel | test_cuda_windows.py | Linux-exported program |
|---|---|---|---|---|---|
| 13.0 | 13.0.88 (installer) | 2.14.1+cu130 | executorch-1.6.0+cu130 | 52 checks pass | runs on the GPU, maxdiff 2.4e-7 |
| 13.2 | 13.2.86 (redist) | 2.14.1+cu132 | executorch-1.6.0+cu132 | 52 checks pass | runs on the GPU, maxdiff 2.4e-7 |
| 13.4 | 13.4.59 (redist) | 2.14.0.dev20260810+cu134 | executorch-1.6.0+cu134 | 52 checks pass | runs on the GPU, maxdiff 2.4e-7 |

- Each wheel built with TORCH_CUDA_ARCH_LIST from cuda_arch_list.sh: device code for exactly
  sm_80/86/89/90/100/120, plus sm_120 PTX. 13.4 runs on a driver that reports 13.2, through
  CUDA's minor-version compatibility.
- Before the /Zc:preprocessor change, 13.2 and 13.4 failed to compile the .cu shims with
  CCCL's "traditional preprocessor" #error; 13.0 was unaffected.
- Negative: the cu134 wheel pointed at the 13.0-exported program fails the train check.
- The export used `test_cuda_windows.py --export` from this branch in WSL, with mingw-w64 and
  the Windows CUDA runtime of the same train (13.0 from install_cuda_windows_cross_compile.sh,
  13.2/13.4 from the redist archives).
- The new CI end-to-end jobs could not be run before upload; they follow cuda-windows.yml's
  linux_job_v3 / windows_job pairing and the steps above.
- Linux: the CMake changes are behind WIN32 / MSVC, or pick the same suffix there (empty on a
  Linux host); setup.py's new path runs only on Windows.
Gasoonjia added a commit that referenced this pull request Oct 3, 2026
Stacked on #23121, which ships linkable libraries in the Windows wheel.

A Windows user can run a CUDA program but cannot lower one: lowering compiles the model
with a toolchain only the Linux side has, and the CUDA Windows CI job already exports on
Linux and runs on Windows for exactly this reason. What was missing is a wheel that gives
that Windows user the runtime. This adds a Windows CUDA wheel carrying the CUDA delegate
as C++ libraries, the same components the Linux CUDA wheel ships (#21645, #21668), and
nothing CUDA in the Python module, which on Windows has no CUDA program to lower. The
export (AOT) side is unchanged and identical to the CPU wheel's; only the runtime links CUDA.

    lib/executorch_backend_cuda.dll    + .lib   the delegate          executorch::backend_cuda
    lib/executorch_extension_cuda.dll  + .lib   the stream helper     executorch::extension_cuda
    backends/cuda/aoti_cuda_shims.dll            the AOTI shim layer   (runtime DLL of the delegate)
    data/lib/aoti_cuda_shims.lib                 already shipped as the lowering stub; also the
                                                 shim layer's import library here
    extension/pybindings/_C.pyd                  unchanged: no CUDA dependency, no CudaBackend

A C++ application links it exactly as on Linux:

    find_package(executorch REQUIRED COMPONENTS kernels_optimized backend_cuda)
    target_link_libraries(app PRIVATE executorch::runtime executorch::kernels_optimized
                                      executorch::backend_cuda)
    # plus the $<TARGET_RUNTIME_DLLS:app> copy from the CPU PR, which now also brings the
    # stream helper and the shim layer

The wheel's DLLs import no CUDA runtime DLL: CMake links cudart.lib into them on Windows,
and that reaches the driver (nvcuda.dll). The compiled model is different: the library
AOTInductor builds into model.pte imports cudart64_13.dll, which comes from the CUDA
Toolkit's bin directory on PATH, as the guide says. PyPI's nvidia-cuda-runtime has win_amd64
builds that carry that DLL, but a C++ program does not search site-packages, so declaring it
would not make it loadable; the wheel declares no NVIDIA package for Windows.

## How

- Generator. The CUDA toolkit's Visual Studio integration fails compiler identification
  under the ClangCL toolset on some toolkit and Visual Studio pairs (MSB4023 in its targets
  file), and the MSVC toolset cannot compile the Python extension's sources. A Windows CUDA
  build therefore uses Ninja Multi-Config with clang-cl, the compiler the CPU wheel uses,
  and nvcc with cl.exe as its host compiler, when ninja is on PATH. Without it, as in a
  --no-build-isolation source install, the build keeps the Visual Studio generator it always
  used. CMAKE_CUDA_COMPILER is passed only when CUDACXX is unset, so options a user put in
  CUDACXX (such as -allow-unsupported-compiler) still apply. The toolkit is the one
  install_utils already reports, so what compiles is the train packaging declares.
- CUDA 13.2 and newer. Their CCCL stops with an #error under cl.exe's traditional
  preprocessor, so the CUDA sources pass -Xcompiler=/Zc:preprocessor on Windows (13.0
  accepts it too).
- flatc / flatcc byproducts and imported locations share one host executable suffix chosen
  by CMAKE_HOST_WIN32. CMAKE_EXECUTABLE_SUFFIX is the target's: a Zephyr toolchain sets
  .elf, and naming the byproduct after it left Ninja with no rule for the host tool.
- The shims resolve the runtime from executorch.dll rather than a static core. (The PAL's
  strong symbols are in #23121; C++20 for the delegate and shims comes from #23135.)
- The shims keep WINDOWS_EXPORT_ALL_SYMBOLS OFF in a static build, as upstream sets it,
  and turn it ON in the shared layout: the delegate DLL imports the stream guard functions
  (get/set/peek/clearCurrentCUDAStream) from the shim DLL, and those carry no
  AOTI_SHIM_EXPORT, so the annotations alone would not export them.
- #23121 keeps a Windows CUDA build static; this drops that guard, so the pybind preset
  now builds the shared layout with CUDA too, and adds the CUDA import libraries to the
  Windows packaging list.
- extension/cuda links CUDA::cudart before backends/cuda finds the toolkit, so the root
  finds it first; a CPU torch does not define that target.
- The package config adds the stream helper and the shim layer to the delegate's runtime
  DLL set, and reports EXECUTORCH_RUNTIME_DLLS_EXTRA for the pre-3.28 route, since the
  shim layer ships outside lib/.

## CI

- build-wheels-cuda-windows.yml mirrors the Linux CUDA workflow on the Windows wheel
  builder and reuses filter_cuda_matrix.py. pre_build_script.sh drops the blanket Windows
  CUDA=OFF (a CPU row is still forced off by the generic row classification), resolves
  TORCH_CUDA_ARCH_LIST from cuda_arch_list.sh, and writes it to BUILD_ENV_FILE as well as
  GITHUB_ENV: the reusable workflow sources that file after the hook, and its own list had
  replaced the row's (sm_75 in, sm_89 out). It also installs ninja for the generator.
- End-to-end, per published CUDA train, the route a Windows user has, modeled on
  cuda-windows.yml. The trains are read from filter_cuda_matrix.py's SUPPORTED_CUDA_VERSIONS
  (today cu132 and cu134; main stopped publishing cu130 in #23380), the same list the release
  rows are built from, so adding or
  dropping a train there changes what is tested. (install_utils.py's list also has 12.6 for
  source builds, but 12.6 is no longer a published wheel train.) A Linux GPU job exports a
  Windows-target program with that train's torch (`test_cuda_windows.py --export`), and a
  Windows GPU job builds the wheel with that train's nvcc, installs it and runs
  test_cuda_windows.py with the program (test_cuda_windows_e2e.ps1). Trains the images do not
  carry are assembled from NVIDIA's checksummed redistributable archives by
  install_cuda_redist.py, which needs no installer or admin rights and fails on a train it
  has no release for. Each train's Windows job waits only for its own program, so one train
  failing to export does not skip the others. The wheel build's own smoke test has no GPU
  program and says so.

## Tests

test_cuda_windows.py, the Windows counterpart of test_cuda_linux.py:

- the CUDA DLLs and import libraries ship;
- _C depends on no CUDA DLL and registers no CUDA backend;
- no shipped CUDA DLL imports cudart64_*.dll (only the driver), and no Windows NVIDIA
  requirement is declared;
- every DLL with device code carries exactly the row's GPUs, both directions, with the
  newest also as PTX; the expected list comes from cuda_arch_list.sh keyed by the wheel's
  +cuXYZ tag, not from the TORCH_CUDA_ARCH_LIST the build consumed;
- a C++ application linking executorch::backend_cuda gets the three DLLs copied beside it
  and sees CudaBackend registered;
- a program lowered on Linux for Windows runs through the C++ SDK and matches eager, after
  checking it was exported with the same CUDA train as the wheel (every CUDA 13 train loads
  cudart64_13.dll, so a mismatch would not fail on its own);
- then what a CPU Windows row checks: test_shared_libraries.py (with the three CUDA
  ownership rows), test_cpp_sdk.py and the MobileNetV3/XNNPACK model run. The Arm Cortex-M
  Python module is not built with Ninja, so its install check is left out.

## Test plan

On Windows 11 x64 with an RTX 5080 (sm_120, driver 595.95), VS 2022 BuildTools, CPython 3.12,
and WSL Ubuntu for the export, once per CUDA train. 13.2 and 13.4 are the published trains
and the ones CI covers; 13.0 is no longer published (#23380) but still builds from source, so
it was run too:

| train | Windows nvcc | WSL export torch | wheel | test_cuda_windows.py | Linux-exported program |
|---|---|---|---|---|---|
| 13.0 | 13.0.88 (installer) | 2.14.1+cu130 | executorch-1.6.0+cu130 | 52 checks pass | runs on the GPU, maxdiff 2.4e-7 |
| 13.2 | 13.2.86 (redist) | 2.14.1+cu132 | executorch-1.6.0+cu132 | 52 checks pass | runs on the GPU, maxdiff 2.4e-7 |
| 13.4 | 13.4.59 (redist) | 2.14.0.dev20260810+cu134 | executorch-1.6.0+cu134 | 52 checks pass | runs on the GPU, maxdiff 2.4e-7 |

- Each wheel built with TORCH_CUDA_ARCH_LIST from cuda_arch_list.sh: device code for exactly
  sm_80/86/89/90/100/120, plus sm_120 PTX. 13.4 runs on a driver that reports 13.2, through
  CUDA's minor-version compatibility.
- Before the /Zc:preprocessor change, 13.2 and 13.4 failed to compile the .cu shims with
  CCCL's "traditional preprocessor" #error; 13.0 was unaffected.
- Negative: the cu134 wheel pointed at the 13.0-exported program fails the train check.
- The export used `test_cuda_windows.py --export` from this branch in WSL, with mingw-w64 and
  the Windows CUDA runtime of the same train (13.0 from install_cuda_windows_cross_compile.sh,
  13.2/13.4 from the redist archives).
- The new CI end-to-end jobs could not be run before upload; they follow cuda-windows.yml's
  linux_job_v3 / windows_job pairing and the steps above.
- Linux: the CMake changes are behind WIN32 / MSVC, or pick the same suffix there (empty on a
  Linux host); setup.py's new path runs only on Windows.

@shoumikhin shoumikhin left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This looks good to me. Two small things inline, neither blocks.

Comment thread tools/cmake/Utils.cmake
set_target_properties(
${target_name} PROPERTIES EXECUTORCH_ANCHOR "${_anchor}"
)
target_link_options(${target_name} INTERFACE "LINKER:/INCLUDE:${_anchor}")

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

minor

I have not run this on Windows, but I expect import executorch.extension.pybindings.data_loader as the first import to fail with a DLL load error now. That module uses no runtime function, yet it gets this /INCLUDE, so it imports executorch_anchor_executorch from executorch.dll. Nothing on its import path adds the lib directory to the DLL search path: its packages have no __init__.py, and it is not one of the four entry points your description lists. The import test imports portable_lib first, so it does not see this. Skipping the anchor for this module, or adding the directory for it, should fix it.

// program built with the other one (/MDd and /MTd define _DEBUG, /MD and /MT do
// not) sees differently laid out types: the two mixed in one process corrupt
// memory instead of failing to link.
#if defined(ET_PREBUILT_RELEASE_CRT) && defined(_DEBUG)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

minor

This checks _DEBUG only. A Release program built with /D_ITERATOR_DEBUG_LEVEL=1 still compiles against the package, and at that level the standard containers carry an extra pointer, so their layout differs from the DLLs'. Checking _ITERATOR_DEBUG_LEVEL too would catch it.

This branch was successfully deployed

1 active deployment
cadence — 813c4764 Deployed Oct 3, 2026 by Gasoonjia via hifi-op-test / hifi4 #31260
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants