Conversation
🔗 Helpful Links🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/23123
Note: Links to docs will display an error until the docs builds have been completed. ❌ 1 Pending, 5 Unclassified FailuresAs of commit dd388dd with merge base 3794e44 ( UNCLASSIFIED FAILURES - DrCI could not classify the following jobs because the workflow did not run on the merge base. The failures may be pre-existing on trunk or introduced by this PR:
This comment was automatically generated by Dr. CI and updates every 15 minutes. |
162d4c7 to
b10b9e3
Compare
5afcafd to
c518096
Compare
b10b9e3 to
886edad
Compare
c518096 to
c2bace9
Compare
886edad to
f5e316b
Compare
24b4561 to
bd4757a
Compare
c2bace9 to
a2310ce
Compare
a2310ce to
5d6a847
Compare
bd4757a to
813c476
Compare
5d6a847 to
cc32b23
Compare
cc32b23 to
87e3b7a
Compare
87e3b7a to
c9578b4
Compare
shoumikhin
left a comment
There was a problem hiding this comment.
Two blockers are left, both inline: the duplicate allocator in the Windows source install, and the end to end job that has no NVIDIA driver. The earlier threads are resolved.
| # carry no AOTI_SHIM_EXPORT, so the annotations alone would not export them; | ||
| # export everything, as every other shipped DLL does. | ||
| set_target_properties( | ||
| aoti_cuda_shims PROPERTIES WINDOWS_EXPORT_ALL_SYMBOLS ON |
There was a problem hiding this comment.
blocker
On Windows with the CUDA toolkit installed, a source install that builds the CUDA backend with the Visual Studio generator no longer links. CUDA is on by default when nvcc is found, and install_executorch uses that generator because nothing installs ninja for it.
runtime/cuda_allocator.cpp is built into aoti_cuda_shims, and on MSVC it is also built into aoti_cuda_backend a few lines below. Now that the Windows build is shared and the shims DLL exports everything, the linker sees CudaAllocator::instance, memcpy_async and release_cached_memory twice and stops with "duplicate symbol".
CI hides it. All 16 test-models-windows and test-model-cuda-windows-e2e jobs hit it, build no wheel, and still end green. The same test-models-windows jobs on #23121 all build the wheel.
The Ninja wheel build links, but the wheel has the same problem in another form: aoti_cuda_shims.dll and executorch_backend_cuda.dll both export the same 13 CudaAllocator symbols, so the process has two allocators.
I think one condition fixes both: add cuda_allocator.cpp to the delegate only for the static layout, so in the shared layout it takes the allocator from the shims DLL, like it already does for the stream functions. I have not built that change.
| uses: pytorch/test-infra/.github/workflows/windows_job.yml@main | ||
| with: | ||
| timeout: 240 | ||
| runner: windows.g5.4xlarge.nvidia.gpu |
There was a problem hiding this comment.
blocker
The end to end jobs fail on both published trains because this job has no NVIDIA driver. The diagnostic you added shows it: nvcuda.dll present: False, driver supports CUDA 0, and cudaGetDeviceCount -> 35, which is cudaErrorInsufficientDriver. That is why forward failed with cudaGetDevice failed.
windows_job.yml has no driver step. The wheel build job does: it runs driver_update.bat. Could this job install the driver before the test, and print nvidia-smi, so a driver problem is easy to spot next time?
Until then no CI job has run a CUDA program from this wheel, so the GPU claims in the description rest on your local runs.
| # Not gated on the export job's result: that is one result for the whole matrix, so a | ||
| # single train failing to export skipped every train here. Each train downloads its own | ||
| # program, and a train whose export failed fails here on the missing artifact. | ||
| if: ${{ !cancelled() && needs.cuda-trains.result == 'success' }} |
There was a problem hiding this comment.
major
!cancelled() removes the usual "needed jobs succeeded" check. On a fork pull request the export job is skipped, but this job still runs and then fails to download an artifact that was never made. cuda-windows.yml has a comment about exactly this trap and gates on the export job's result.
Adding the same fork condition here, or needs.export-cuda-windows-wheel-program.result != 'skipped', keeps the "one train failed, run the others" behavior and fixes forks. This pull request is not from a fork, so its own CI cannot show it.
| package-name: executorch | ||
| name: ${{ matrix.repository }} | ||
| uses: pytorch/test-infra/.github/workflows/build_wheels_windows.yml@main | ||
| with: |
There was a problem hiding this comment.
major
The Windows CUDA wheel build is too close to its time limit. This job sets no timeout, so it gets the reusable workflow's default of 60 minutes. The same job took 58 minutes on an earlier push and hit the 60 minute limit on the latest one, in the smoke test, so no wheel was uploaded. On the latest push the smoke test started 55 minutes in, and it took 9 minutes the time before.
Could you set timeout: 90 here, like the export job does?
| ``` | ||
|
|
||
| That export also needs the MinGW cross compiler and the Windows CUDA runtime on the Linux side; | ||
| `.ci/docker/common/install_cuda_windows_cross_compile.sh` installs both. |
There was a problem hiding this comment.
major
install_cuda_windows_cross_compile.sh only knows CUDA 12.6, 12.8, 12.9 and 13.0. For cu132 or cu134 it exits with "CUDA version 13.2 is not in the known version map", and the same for 13.4. It is also a script for building the CI image (it runs apt-get as root).
Could the guide describe the setup the workflow uses instead: the MinGW compiler, plus the Windows CUDA runtime for the same CUDA version, with WINDOWS_CUDA_HOME pointing at it?
| # The Python side of this wheel carries no CUDA, and install_requirements installs CPU | ||
| # torch on Windows, which is what the wheel build and its import checks need. It also | ||
| # installs the torchao nightly the wheel declares, which PyPI does not carry. | ||
| Invoke-Native { python install_requirements.py } |
There was a problem hiding this comment.
major
I expect the last step of test_cuda_windows.py (the MobileNetV3 run) to fail here once the driver is fixed. install_requirements.py without --example does not install torchvision, and the MobileNetV3 model imports it. The wheel does not depend on it either. I have not seen it fail, because the job stops earlier, at the GPU check.
pre_build_script.sh calls install_requirements.sh --example for the same reason. Could this script do the same?
| f"-DCUDAToolkit_ROOT={windows_cuda_home}", | ||
| # Ninja does not generate the Arm Cortex-M Python module's two | ||
| # identically named sources as distinct rules. | ||
| "-DEXECUTORCH_BUILD_CMSIS_NN_PYBINDS=OFF", |
There was a problem hiding this comment.
minor
This drops the Cortex-M Python module (cmsis_nn) from the CUDA wheel, so Cortex-M lowering cannot work from it. The comment says the module has two identically named sources. I think the real clash is that CMSIS-NN's static library and its test-only shared library both write cmsis-nn.lib. A small test project shows that Ninja error with the two libraries, and setting ARCHIVE_OUTPUT_NAME on the shared one avoids it, so it might make this switch unnecessary. I did not try a full build.
The summary says the export side is identical to the CPU wheel's. Without this module it is not, so it would help to say so in the summary too, not only in the Tests section.
Stacked on #23121, which ships linkable libraries in the Windows wheel. A Windows user can run a CUDA program but cannot lower one: lowering compiles the model with a toolchain only the Linux side has, and the CUDA Windows CI job already exports on Linux and runs on Windows for exactly this reason. What was missing is a wheel that gives that Windows user the runtime. This adds a Windows CUDA wheel carrying the CUDA delegate as C++ libraries, the same components the Linux CUDA wheel ships (#21645, #21668), and nothing CUDA in the Python module, which on Windows has no CUDA program to lower. The export (AOT) side is unchanged and identical to the CPU wheel's; only the runtime links CUDA. lib/executorch_backend_cuda.dll + .lib the delegate executorch::backend_cuda lib/executorch_extension_cuda.dll + .lib the stream helper executorch::extension_cuda backends/cuda/aoti_cuda_shims.dll the AOTI shim layer (runtime DLL of the delegate) data/lib/aoti_cuda_shims.lib already shipped as the lowering stub; also the shim layer's import library here extension/pybindings/_C.pyd unchanged: no CUDA dependency, no CudaBackend A C++ application links it exactly as on Linux: find_package(executorch REQUIRED COMPONENTS kernels_optimized backend_cuda) target_link_libraries(app PRIVATE executorch::runtime executorch::kernels_optimized executorch::backend_cuda) # plus the $<TARGET_RUNTIME_DLLS:app> copy from the CPU PR, which now also brings the # stream helper and the shim layer The wheel's DLLs import no CUDA runtime DLL: CMake links cudart.lib into them on Windows, and that reaches the driver (nvcuda.dll). The compiled model is different: the library AOTInductor builds into model.pte imports cudart64_13.dll, which comes from the CUDA Toolkit's bin directory on PATH, as the guide says. PyPI's nvidia-cuda-runtime has win_amd64 builds that carry that DLL, but a C++ program does not search site-packages, so declaring it would not make it loadable; the wheel declares no NVIDIA package for Windows. ## How - Generator. The CUDA toolkit's Visual Studio integration fails compiler identification under the ClangCL toolset on some toolkit and Visual Studio pairs (MSB4023 in its targets file), and the MSVC toolset cannot compile the Python extension's sources. A Windows CUDA build therefore uses Ninja Multi-Config with clang-cl, the compiler the CPU wheel uses, and nvcc with cl.exe as its host compiler, when ninja is on PATH. Without it, as in a --no-build-isolation source install, the build keeps the Visual Studio generator it always used. CMAKE_CUDA_COMPILER is passed only when CUDACXX is unset, so options a user put in CUDACXX (such as -allow-unsupported-compiler) still apply. The toolkit is the one install_utils already reports, so what compiles is the train packaging declares. - CUDA 13.2 and newer. Their CCCL stops with an #error under cl.exe's traditional preprocessor, so the CUDA sources pass -Xcompiler=/Zc:preprocessor on Windows (13.0 accepts it too). - flatc / flatcc byproducts and imported locations share one host executable suffix chosen by CMAKE_HOST_WIN32. CMAKE_EXECUTABLE_SUFFIX is the target's: a Zephyr toolchain sets .elf, and naming the byproduct after it left Ninja with no rule for the host tool. - The shims resolve the runtime from executorch.dll rather than a static core. (The PAL's strong symbols are in #23121; C++20 for the delegate and shims comes from #23135.) - The shims keep WINDOWS_EXPORT_ALL_SYMBOLS OFF in a static build, as upstream sets it, and turn it ON in the shared layout: the delegate DLL imports the stream guard functions (get/set/peek/clearCurrentCUDAStream) from the shim DLL, and those carry no AOTI_SHIM_EXPORT, so the annotations alone would not export them. - #23121 keeps a Windows CUDA build static; this drops that guard, so the pybind preset now builds the shared layout with CUDA too, and adds the CUDA import libraries to the Windows packaging list. - extension/cuda links CUDA::cudart before backends/cuda finds the toolkit, so the root finds it first; a CPU torch does not define that target. - The package config adds the stream helper and the shim layer to the delegate's runtime DLL set, and reports EXECUTORCH_RUNTIME_DLLS_EXTRA for the pre-3.28 route, since the shim layer ships outside lib/. ## CI - build-wheels-cuda-windows.yml mirrors the Linux CUDA workflow on the Windows wheel builder and reuses filter_cuda_matrix.py. pre_build_script.sh drops the blanket Windows CUDA=OFF (a CPU row is still forced off by the generic row classification), resolves TORCH_CUDA_ARCH_LIST from cuda_arch_list.sh, and writes it to BUILD_ENV_FILE as well as GITHUB_ENV: the reusable workflow sources that file after the hook, and its own list had replaced the row's (sm_75 in, sm_89 out). It also installs ninja for the generator. - End-to-end, per published CUDA train, the route a Windows user has, modeled on cuda-windows.yml. The trains are read from filter_cuda_matrix.py's SUPPORTED_CUDA_VERSIONS (today cu132 and cu134; main stopped publishing cu130 in #23380), the same list the release rows are built from, so adding or dropping a train there changes what is tested. (install_utils.py's list also has 12.6 for source builds, but 12.6 is no longer a published wheel train.) A Linux GPU job exports a Windows-target program with that train's torch (`test_cuda_windows.py --export`), and a Windows GPU job builds the wheel with that train's nvcc, installs it and runs test_cuda_windows.py with the program (test_cuda_windows_e2e.ps1). Trains the images do not carry are assembled from NVIDIA's checksummed redistributable archives by install_cuda_redist.py, which needs no installer or admin rights and fails on a train it has no release for. Each train's Windows job waits only for its own program, so one train failing to export does not skip the others. The Windows GPU runners come without an NVIDIA driver (nvcuda.dll is missing and cudaGetDeviceCount returns cudaErrorInsufficientDriver), so the job first installs the same data-center driver, from the same place, that pytorch/pytorch's Windows CUDA smoke tests install (driver_update.bat, 580.88); cu132 and cu134 run on it through CUDA's minor-version compatibility, as PyTorch's own wheels do. The wheel build's own smoke test has no GPU program and says so. ## Tests test_cuda_windows.py, the Windows counterpart of test_cuda_linux.py: - the CUDA DLLs and import libraries ship; - _C depends on no CUDA DLL and registers no CUDA backend; - no shipped CUDA DLL imports cudart64_*.dll (only the driver), and no Windows NVIDIA requirement is declared; - every DLL with device code carries exactly the row's GPUs, both directions, with the newest also as PTX; the expected list comes from cuda_arch_list.sh keyed by the wheel's +cuXYZ tag, not from the TORCH_CUDA_ARCH_LIST the build consumed; - a C++ application linking executorch::backend_cuda gets the three DLLs copied beside it and sees CudaBackend registered; - a program lowered on Linux for Windows runs through the C++ SDK and matches eager, after checking it was exported with the same CUDA train as the wheel (every CUDA 13 train loads cudart64_13.dll, so a mismatch would not fail on its own); - then what a CPU Windows row checks: test_shared_libraries.py (with the three CUDA ownership rows), test_cpp_sdk.py and the MobileNetV3/XNNPACK model run. The Arm Cortex-M Python module is not built with Ninja, so its install check is left out. ## Test plan On Windows 11 x64 with an RTX 5080 (sm_120, driver 595.95), VS 2022 BuildTools, CPython 3.12, and WSL Ubuntu for the export, once per CUDA train. 13.2 and 13.4 are the published trains and the ones CI covers; 13.0 is no longer published (#23380) but still builds from source, so it was run too: | train | Windows nvcc | WSL export torch | wheel | test_cuda_windows.py | Linux-exported program | |---|---|---|---|---|---| | 13.0 | 13.0.88 (installer) | 2.14.1+cu130 | executorch-1.6.0+cu130 | 52 checks pass | runs on the GPU, maxdiff 2.4e-7 | | 13.2 | 13.2.86 (redist) | 2.14.1+cu132 | executorch-1.6.0+cu132 | 52 checks pass | runs on the GPU, maxdiff 2.4e-7 | | 13.4 | 13.4.59 (redist) | 2.14.0.dev20260810+cu134 | executorch-1.6.0+cu134 | 52 checks pass | runs on the GPU, maxdiff 2.4e-7 | - Each wheel built with TORCH_CUDA_ARCH_LIST from cuda_arch_list.sh: device code for exactly sm_80/86/89/90/100/120, plus sm_120 PTX. 13.4 runs on a driver that reports 13.2, through CUDA's minor-version compatibility. - Before the /Zc:preprocessor change, 13.2 and 13.4 failed to compile the .cu shims with CCCL's "traditional preprocessor" #error; 13.0 was unaffected. - Negative: the cu134 wheel pointed at the 13.0-exported program fails the train check. - The export used `test_cuda_windows.py --export` from this branch in WSL, with mingw-w64 and the Windows CUDA runtime of the same train (13.0 from install_cuda_windows_cross_compile.sh, 13.2/13.4 from the redist archives). - The new CI end-to-end jobs could not be run before upload; they follow cuda-windows.yml's linux_job_v3 / windows_job pairing and the steps above. - Linux: the CMake changes are behind WIN32 / MSVC, or pick the same suffix there (empty on a Linux host); setup.py's new path runs only on Windows.
c9578b4 to
dd388dd
Compare
Stacked on #23121, which ships linkable libraries in the Windows wheel.
A Windows user can run a CUDA program but cannot lower one: lowering compiles the model
with a toolchain only the Linux side has, and the CUDA Windows CI job already exports on
Linux and runs on Windows for exactly this reason. What was missing is a wheel that gives
that Windows user the runtime. This adds a Windows CUDA wheel carrying the CUDA delegate
as C++ libraries, the same components the Linux CUDA wheel ships (#21645, #21668), and
nothing CUDA in the Python module, which on Windows has no CUDA program to lower. The
export (AOT) side is unchanged and identical to the CPU wheel's; only the runtime links CUDA.
A C++ application links it exactly as on Linux:
The wheel's DLLs import no CUDA runtime DLL: CMake links cudart.lib into them on Windows,
and that reaches the driver (nvcuda.dll). The compiled model is different: the library
AOTInductor builds into model.pte imports cudart64_13.dll, which comes from the CUDA
Toolkit's bin directory on PATH, as the guide says. PyPI's nvidia-cuda-runtime has win_amd64
builds that carry that DLL, but a C++ program does not search site-packages, so declaring it
would not make it loadable; the wheel declares no NVIDIA package for Windows.
How
under the ClangCL toolset on some toolkit and Visual Studio pairs (MSB4023 in its targets
file), and the MSVC toolset cannot compile the Python extension's sources. A Windows CUDA
build therefore uses Ninja Multi-Config with clang-cl, the compiler the CPU wheel uses,
and nvcc with cl.exe as its host compiler, when ninja is on PATH. Without it, as in a
--no-build-isolation source install, the build keeps the Visual Studio generator it always
used. CMAKE_CUDA_COMPILER is passed only when CUDACXX is unset, so options a user put in
CUDACXX (such as -allow-unsupported-compiler) still apply. The toolkit is the one
install_utils already reports, so what compiles is the train packaging declares.
preprocessor, so the CUDA sources pass -Xcompiler=/Zc:preprocessor on Windows (13.0
accepts it too).
by CMAKE_HOST_WIN32. CMAKE_EXECUTABLE_SUFFIX is the target's: a Zephyr toolchain sets
.elf, and naming the byproduct after it left Ninja with no rule for the host tool.
strong symbols are in Ship linkable libraries in the Windows CPU wheel #23121; C++20 for the delegate and shims comes from Fix CUDA backend Windows build under C++20 #23135.)
and turn it ON in the shared layout: the delegate DLL imports the stream guard functions
(get/set/peek/clearCurrentCUDAStream) from the shim DLL, and those carry no
AOTI_SHIM_EXPORT, so the annotations alone would not export them.
now builds the shared layout with CUDA too, and adds the CUDA import libraries to the
Windows packaging list.
finds it first; a CPU torch does not define that target.
DLL set, and reports EXECUTORCH_RUNTIME_DLLS_EXTRA for the pre-3.28 route, since the
shim layer ships outside lib/.
CI
builder and reuses filter_cuda_matrix.py. pre_build_script.sh drops the blanket Windows
CUDA=OFF (a CPU row is still forced off by the generic row classification), resolves
TORCH_CUDA_ARCH_LIST from cuda_arch_list.sh, and writes it to BUILD_ENV_FILE as well as
GITHUB_ENV: the reusable workflow sources that file after the hook, and its own list had
replaced the row's (sm_75 in, sm_89 out). It also installs ninja for the generator.
cuda-windows.yml. The trains are read from filter_cuda_matrix.py's SUPPORTED_CUDA_VERSIONS
(today cu132 and cu134; main stopped publishing cu130 in Stop publishing CUDA 13.0 wheels #23380), the same list the release
rows are built from, so adding or
dropping a train there changes what is tested. (install_utils.py's list also has 12.6 for
source builds, but 12.6 is no longer a published wheel train.) A Linux GPU job exports a
Windows-target program with that train's torch (
test_cuda_windows.py --export), and aWindows GPU job builds the wheel with that train's nvcc, installs it and runs
test_cuda_windows.py with the program (test_cuda_windows_e2e.ps1). Trains the images do not
carry are assembled from NVIDIA's checksummed redistributable archives by
install_cuda_redist.py, which needs no installer or admin rights and fails on a train it
has no release for. Each train's Windows job waits only for its own program, so one train
failing to export does not skip the others. The Windows GPU runners come without an NVIDIA
driver (nvcuda.dll is missing and cudaGetDeviceCount returns cudaErrorInsufficientDriver),
so the job first installs the same data-center driver, from the same place, that
pytorch/pytorch's Windows CUDA smoke tests install (driver_update.bat, 580.88); cu132 and
cu134 run on it through CUDA's minor-version compatibility, as PyTorch's own wheels do. The
wheel build's own smoke test has no GPU program and says so.
Tests
test_cuda_windows.py, the Windows counterpart of test_cuda_linux.py:
requirement is declared;
newest also as PTX; the expected list comes from cuda_arch_list.sh keyed by the wheel's
+cuXYZ tag, not from the TORCH_CUDA_ARCH_LIST the build consumed;
and sees CudaBackend registered;
checking it was exported with the same CUDA train as the wheel (every CUDA 13 train loads
cudart64_13.dll, so a mismatch would not fail on its own);
ownership rows), test_cpp_sdk.py and the MobileNetV3/XNNPACK model run. The Arm Cortex-M
Python module is not built with Ninja, so its install check is left out.
Test plan
On Windows 11 x64 with an RTX 5080 (sm_120, driver 595.95), VS 2022 BuildTools, CPython 3.12,
and WSL Ubuntu for the export, once per CUDA train. 13.2 and 13.4 are the published trains
and the ones CI covers; 13.0 is no longer published (#23380) but still builds from source, so
it was run too:
sm_80/86/89/90/100/120, plus sm_120 PTX. 13.4 runs on a driver that reports 13.2, through
CUDA's minor-version compatibility.
CCCL's "traditional preprocessor" #error; 13.0 was unaffected.
test_cuda_windows.py --exportfrom this branch in WSL, with mingw-w64 andthe Windows CUDA runtime of the same train (13.0 from install_cuda_windows_cross_compile.sh,
13.2/13.4 from the redist archives).
linux_job_v3 / windows_job pairing and the steps above.
Linux host); setup.py's new path runs only on Windows.