Ship a pkg-config file for the runtime in the Linux and macOS wheels - #23169
Merged
Merged
Conversation
The wheel ships the runtime as shared libraries with headers and a CMake package, but build systems that read pkg-config, such as Meson, cannot use it. They look for executorch.pc and the wheel has none, so today they need a full source build and install just to get that one file. This adds lib/pkgconfig/executorch.pc to the wheel. It works out every path from its own location, so the wheel can be installed anywhere. It passes the same compile definitions as the CMake package, including the event tracer switch when the libraries are built with it, and it adds a runtime search path so a program finds the libraries without LD_LIBRARY_PATH. Like the CMake runtime component, it covers the runtime only, and the kernel libraries are named on the link line. Windows is left out because the Windows wheel does not ship the runtime library.
shoumikhin
requested review from
kirklandsign,
larryliu0820 and
mergennachin
as code owners
September 25, 2026 21:48
🔗 Helpful Links🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/23169
Note: Links to docs will display an error until the docs builds have been completed. ✅ No FailuresAs of commit fd6587f with merge base 4199643 ( This comment was automatically generated by Dr. CI and updates every 15 minutes. |
The pkg-config file left out the thread pool. The CMake package adds ET_USE_THREADPOOL and the thread pool library to the runtime whenever the wheel ships the thread pool. Without them, code built from the pkg-config flags gets the serial fallback of parallel_for, with no error. The file now adds both under the same condition. On Linux the runtime search path was joined to --enable-new-dtags in one linker argument. Meson only keeps a dependency's search path when the argument starts with -Wl,-rpath, so meson install removed it and the installed program could not find the libraries. They are now two arguments. The smoke test now also checks that the compile definitions pkg-config prints match the ones the CMake package gives, so the two cannot drift. The docs now show separate Linux and macOS commands, scope --no-as-needed with push-state and pop-state, add to PKG_CONFIG_PATH instead of replacing it, ask for an absolute path, and give a Meson example that was built, installed and run.
JakeStevens
approved these changes
Sep 26, 2026
Gasoonjia
added a commit
that referenced
this pull request
Sep 29, 2026
The Windows wheel shipped one fused Python extension and no linkable libraries, so a C++ application got nothing from it: the headers and the CMake package were there, but every component a consumer asked for resolved to nothing. Linux and macOS already ship the runtime, the kernels, the delegates and the thread pool as separate libraries that a C++ application links directly (#21610, #21771). This does the same for Windows. Before (Windows x64, CPython 3.12): extension/pybindings/_C.cp312-win_amd64.pyd 8.27 MB <- everything fused in here lib/ does not exist After: lib/executorch_kernels_optimized.dll 5.02 MB lib/executorch_backend_xnnpack.dll 2.47 MB lib/executorch.dll 0.38 MB lib/executorch_kernels_quantized.dll 0.22 MB lib/executorch_threadpool.dll 0.21 MB lib/executorch_etdump.dll 0.05 MB (loaded by _C; not offered to C++) lib/*.lib <- import libraries a consumer links extension/pybindings/_C.cp312-win_amd64.pyd 0.63 MB <- just the bindings now ## How The three mechanisms that made the shared layout Linux and macOS only now have Windows equivalents: what ELF / Mach-O Windows exporting the runtime API default visibility WINDOWS_EXPORT_ALL_SYMBOLS keeping a registration-only lib --no-as-needed / named dylib /INCLUDE of a per-DLL anchor finding sibling libraries $ORIGIN / @loader_path os.add_dll_directory, and a consumer copies the DLLs - Exports. The runtime has no dllexport annotations, which is why CMakeLists.txt refused a shared Windows build. Each shipped component DLL now exports all of its symbols, unless its target already exports through its own annotations (WINDOWS_EXPORT_ALL_SYMBOLS OFF), which it keeps. CMake builds that list from a target's own objects and cannot read an archive, so the runtime DLL takes its components as objects rather than through /WHOLEARCHIVE. That uses $<COMPILE_ONLY>, new in CMake 3.27, so a shared Windows build asks for 3.27 with a clear message, and the Windows build requirement is raised to match. - Retention. A PE import survives only if some symbol from it is referenced, so each shipped DLL exports `executorch_anchor_<name>` and the imported targets carry `/INCLUDE` of it, the counterpart of the Linux `--no-as-needed`. - One registry. A consumer names the runtime's import library ahead of any archive, as on Linux, so nothing resolves the registry from a private static copy. Without this the optimized kernels registered into their own table: 28 operators visible instead of 242. - Loading. A DLL records no search path. The Python entry points that load a DLL depending on executorch/lib register that directory first (portable_lib, kernels.quantized, codegen.tools, llm.custom_ops.op_tile_crop_aot), computed without resolving symlinks so an editable install, where those directories are links, registers its own lib too. A C++ consumer copies the DLLs beside its executable with `$<TARGET_RUNTIME_DLLS>`, which the C++ guide now shows, including in the quick start. - Debug. The DLLs use the release C++ library, and a Debug consumer mixing in the debug one would corrupt memory. The package defines ET_PREBUILT_RELEASE_CRT for its consumers, and the runtime headers refuse a Debug build with it at compile time, with a message asking for Release. (A /FAILIFMISMATCH link option would not do: link.exe ignores it on the command line and honours it only inside an object file.) CUDA is not part of this. A Windows build with CUDA on keeps the static layout it had, since the Windows CUDA delegate has not been built as a DLL; a follow-up stacked on this ships it. Two bugs only the shared layout exposes on Windows are fixed here: - The process hung at exit about half the time. Windows terminates worker threads before running a DLL's static destructors, so destroying the global thread pool there waited on a lock a terminated worker could hold (`LdrShutdownProcess -> pthreadpool_destroy -> mtx_lock`). The pool is now deliberately not destroyed on Windows; Linux and macOS are unchanged. - `quantized_ops_aot_lib.dll` failed to load, which `executorch.kernels.quantized` swallows, so quantized export silently lost its out variants. The Windows wheel is built without the event tracer (unchanged), so the `etdump` component is not offered to C++ there rather than handing a consumer a profiler that records nothing. The pkg-config file the Linux and macOS wheels ship (#23169) is not generated on Windows either, since its link flags are GNU-style and a DLL records no search path for them; the test checks that it is absent. The pre-3.28 variables route links the import libraries and states C++20, which the Windows runtime headers need. ## Tests The two wheel suites now run on Windows, as they do on macOS since #21771, wired into test_windows.py. Their Windows forms ask the platform's own questions: - symbols: dumpbin /exports matched against MSVC decorated names, and for code a DLL takes from an archive, which the export list cannot name, the import direction: the owner must import `register_kernels` / `register_backend` from executorch.dll; - dependencies: dumpbin /dependents and /imports in place of readelf / otool; - loading: LoadLibrary per binary in its own process, which binds every import as ldd -r does, from the package and from a relocated copy; - paths: every recorded dependency is a bare DLL name; - platform tag: the PE machine field of every binary against win_amd64; - C++ consumers: built with --config Release, DLLs copied beside the executable, and the installed package removed from PATH so nothing resolves it for them; - package entry points: the extensions are imported the way a user reaches them (the pybindings through portable_lib), with nothing added to the DLL search path, and `import executorch.kernels.quantized` has to register the quantized out variants, so the entry points' own registration is what is tested; - a Debug consumer has to be refused; an application using the thread pool has to exit five runs in a row; every program and probe the suites run has a timeout, so a hang fails as a named timeout rather than the whole job timing out. ## Test plan On Windows 11 x64, VS 2022 BuildTools with ClangCL, CPython 3.12, from a wheel built with `python setup.py bdist_wheel` and installed into a clean environment: - `.ci/scripts/wheel/test_windows.py`: passes, including every check in test_shared_libraries.py and test_cpp_sdk.py (single owner of each component, _C contains no component, every DLL loads in place and relocated, custom op registers, platform tag; version request, 121 of 123 headers compile, documented example builds, runtime alone has no kernels, kernels / quantized / XNNPACK consumers match eager PyTorch, delegate fails without its component, pre-3.28 route on CMake 3.24, relocation, one shared registry) and MobileNetV3 through XNNPACK matching eager. - The exit hang: 8 of 20 and 10 of 20 runs of two consumers hung before the thread pool fix, 0 of 20 after. - Linux and macOS: the CMake changes are behind WIN32 / MSVC and the Python changes behind sys.platform == "win32"; relying on their wheel rows for confirmation.
Gasoonjia
added a commit
that referenced
this pull request
Sep 30, 2026
The Windows wheel shipped one fused Python extension and no linkable libraries, so a C++ application got nothing from it: the headers and the CMake package were there, but every component a consumer asked for resolved to nothing. Linux and macOS already ship the runtime, the kernels, the delegates and the thread pool as separate libraries that a C++ application links directly (#21610, #21771). This does the same for Windows. Before (Windows x64, CPython 3.12): extension/pybindings/_C.cp312-win_amd64.pyd 8.27 MB <- everything fused in here lib/ does not exist After: lib/executorch_kernels_optimized.dll 5.02 MB lib/executorch_backend_xnnpack.dll 2.47 MB lib/executorch.dll 0.38 MB lib/executorch_kernels_quantized.dll 0.22 MB lib/executorch_threadpool.dll 0.21 MB lib/executorch_etdump.dll 0.05 MB (loaded by _C; not offered to C++) lib/*.lib <- import libraries a consumer links extension/pybindings/_C.cp312-win_amd64.pyd 0.63 MB <- just the bindings now ## How The three mechanisms that made the shared layout Linux and macOS only now have Windows equivalents: what ELF / Mach-O Windows exporting the runtime API default visibility WINDOWS_EXPORT_ALL_SYMBOLS keeping a registration-only lib --no-as-needed / named dylib /INCLUDE of a per-DLL anchor finding sibling libraries $ORIGIN / @loader_path os.add_dll_directory, and a consumer copies the DLLs - Exports. The runtime has no dllexport annotations, which is why CMakeLists.txt refused a shared Windows build. Each shipped component DLL now exports all of its symbols, unless its target already exports through its own annotations (WINDOWS_EXPORT_ALL_SYMBOLS OFF), which it keeps. CMake builds that list from a target's own objects and cannot read an archive, so the runtime DLL takes its components as objects rather than through /WHOLEARCHIVE. The caller says which archives are part of a DLL's exports (the helper's EXPORTED option) instead of the helper inferring it from a property another call sets, so the result does not depend on call order. That uses $<COMPILE_ONLY>, new in CMake 3.27, so a shared Windows build asks for 3.27 with a clear message, the Windows build requirement is raised to match, and executorch_shared publishes the cxx_std_20 its headers need on Windows. - Retention. A PE import survives only if some symbol from it is referenced, so each shipped DLL exports `executorch_anchor_<name>` and the imported targets carry `/INCLUDE` of it, the counterpart of the Linux `--no-as-needed`. The anchor source is generated per configuration, so a multi-config build with CMAKE_DEBUG_POSTFIX works. - One registry. A consumer names the runtime's import library ahead of any archive, as on Linux, so nothing resolves the registry from a private static copy. Without this the optimized kernels registered into their own table: 28 operators visible instead of 242. - Loading. A DLL records no search path. The Python entry points that load a DLL depending on executorch/lib register that directory first (portable_lib, kernels.quantized, codegen.tools, llm.custom_ops.op_tile_crop_aot), computed without resolving symlinks so an editable install, where those directories are links, registers its own lib too. A C++ consumer copies the DLLs beside its executable with `$<TARGET_RUNTIME_DLLS>`, which the C++ guide now shows, including in the quick start. - Debug. The DLLs use the C++ library of the configuration the wheel was built in, and a consumer in the other configuration would mix the two and corrupt memory. The package defines ET_PREBUILT_RELEASE_CRT, or ET_PREBUILT_DEBUG_CRT for a DEBUG=1 build, and the runtime headers refuse the mismatched configuration at compile time with a message naming the right one. (A /FAILIFMISMATCH link option would not do: link.exe ignores it on the command line and honours it only inside an object file.) - cpuinfo. The thread pool DLL exports pthreadpool, so the process has one pool, but keeps cpuinfo private. cpuinfo's feature checks (cpuinfo_has_x86_avx2 and the rest) are inline reads of cpuinfo_isa, which cpuinfo.h does not declare dllimport, so a DLL can only read its own copy; with cpuinfo exported, XNNPACK initialized the thread pool's copy and read its own, all zeros, and ran baseline kernels. Each DLL that uses cpuinfo now carries all of it. Measured on an x64 machine with AVX2: an XNNPACK-delegated 2 x Linear(1024) model at batch 256 went from 2.05-2.14 ms to 1.47-1.52 ms per run, with the same results. CUDA is not part of this. A Windows build with CUDA on keeps the static layout it had, since the Windows CUDA delegate has not been built as a DLL; a follow-up stacked on this ships it. Two bugs only the shared layout exposes on Windows are fixed here: - The process hung at exit about half the time. Windows terminates worker threads before running a DLL's static destructors, so destroying the global thread pool there waited on a lock a terminated worker could hold (`LdrShutdownProcess -> pthreadpool_destroy -> mtx_lock`). The pool is now deliberately not destroyed on Windows; Linux and macOS are unchanged. - `quantized_ops_aot_lib.dll` failed to load, which `executorch.kernels.quantized` swallows, so quantized export silently lost its out variants. The Windows wheel is built without the event tracer (unchanged), so the `etdump` component is not offered to C++ there rather than handing a consumer a profiler that records nothing. The pkg-config file the Linux and macOS wheels ship (#23169) is not generated on Windows either, since its link flags are GNU-style and a DLL records no search path for them; the test checks that it is absent. The pre-3.28 variables route links the import libraries and states C++20, which the Windows runtime headers need. ## Tests The two wheel suites now run on Windows, as they do on macOS since #21771, wired into test_windows.py. Their Windows forms ask the platform's own questions: - symbols: dumpbin /exports matched against MSVC decorated names, and for code a DLL takes from an archive, which the export list cannot name, the import direction: the owner must import `register_kernels` / `register_backend` from executorch.dll; - dependencies: dumpbin /dependents and /imports in place of readelf / otool; - loading: LoadLibrary per binary in its own process, which binds every import as ldd -r does, from the package and from a relocated copy; - paths: every recorded dependency is a bare DLL name; - platform tag: the PE machine field of every binary against win_amd64; - C++ consumers: built with --config Release, DLLs copied beside the executable, and the installed package removed from PATH so nothing resolves it for them; - package entry points: the extensions are imported the way a user reaches them (the pybindings through portable_lib), with nothing added to the DLL search path, and `import executorch.kernels.quantized` has to register the quantized out variants, so the entry points' own registration is what is tested; - a consumer in the wrong configuration has to be refused; an application using the thread pool has to exit five runs in a row; no DLL may import cpuinfo from another (the split state has no numeric symptom, only slower kernels); every program the suites build and every Python probe or export they start has a timeout, so a hang fails as a named timeout rather than the whole job timing out. Compiler, CMake and binary tool invocations are not bounded. ## Test plan On Windows 11 x64, VS 2022 BuildTools with ClangCL, CPython 3.12, from a wheel built with `python setup.py bdist_wheel` and installed into a clean environment: - `.ci/scripts/wheel/test_windows.py`: passes, including every check in test_shared_libraries.py and test_cpp_sdk.py (single owner of each component, _C contains no component, every DLL loads in place and relocated, custom op registers, platform tag; version request, 121 of 123 headers compile, documented example builds, runtime alone has no kernels, kernels / quantized / XNNPACK consumers match eager PyTorch, delegate fails without its component, pre-3.28 route on CMake 3.24, relocation, one shared registry) and MobileNetV3 through XNNPACK matching eager. - The exit hang: 8 of 20 and 10 of 20 runs of two consumers hung before the thread pool fix, 0 of 20 after. - Linux and macOS: the CMake changes are behind WIN32 / MSVC and the Python changes behind sys.platform == "win32"; relying on their wheel rows for confirmation.
Gasoonjia
added a commit
that referenced
this pull request
Oct 2, 2026
The Windows wheel shipped one fused Python extension and no linkable libraries, so a C++ application got nothing from it: the headers and the CMake package were there, but every component a consumer asked for resolved to nothing. Linux and macOS already ship the runtime, the kernels, the delegates and the thread pool as separate libraries that a C++ application links directly (#21610, #21771). This does the same for Windows. Before (Windows x64, CPython 3.12): extension/pybindings/_C.cp312-win_amd64.pyd 8.27 MB <- everything fused in here lib/ does not exist After: lib/executorch_kernels_optimized.dll 5.02 MB lib/executorch_backend_xnnpack.dll 2.47 MB lib/executorch.dll 0.38 MB lib/executorch_kernels_quantized.dll 0.22 MB lib/executorch_threadpool.dll 0.21 MB lib/executorch_etdump.dll 0.05 MB (loaded by _C; not offered to C++) lib/*.lib <- import libraries a consumer links extension/pybindings/_C.cp312-win_amd64.pyd 0.63 MB <- just the bindings now ## How The three mechanisms that made the shared layout Linux and macOS only now have Windows equivalents: what ELF / Mach-O Windows exporting the runtime API default visibility WINDOWS_EXPORT_ALL_SYMBOLS keeping a registration-only lib --no-as-needed / named dylib /INCLUDE of a per-DLL anchor finding sibling libraries $ORIGIN / @loader_path os.add_dll_directory, and a consumer copies the DLLs - Exports. The runtime has no dllexport annotations, which is why CMakeLists.txt refused a shared Windows build. Each shipped component DLL now exports all of its symbols, unless its target already exports through its own annotations (WINDOWS_EXPORT_ALL_SYMBOLS OFF), which it keeps. CMake builds that list from a target's own objects and cannot read an archive, so the runtime DLL takes its components as objects rather than through /WHOLEARCHIVE. The caller says which archives are part of a DLL's exports (the helper's EXPORTED option) instead of the helper inferring it from a property another call sets, so the result does not depend on call order. That uses $<COMPILE_ONLY>, new in CMake 3.27, so a shared Windows build asks for 3.27 with a clear message, the Windows build requirement is raised to match, and executorch_shared publishes the cxx_std_20 its headers need on Windows. The PAL source is built with its C functions strong there: the export list skips weak symbols, so the DLL exported none of the et_pal_* functions and clock.h's inline ticks_to_ns() failed to link against it. - Retention. A PE import survives only if some symbol from it is referenced, so each shipped DLL exports `executorch_anchor_<name>` and the imported targets carry `/INCLUDE` of it, the counterpart of the Linux `--no-as-needed`. The name follows the shipped file name without any configuration postfix, and is resolved at the end of configure and written out literally, so it survives `cmake --install` (a generator expression would be evaluated against the imported target, which has no OUTPUT_NAME) and is the same in every configuration of a multi-config build. - One registry. A consumer names the runtime's import library ahead of any archive, as on Linux, so nothing resolves the registry from a private static copy. Without this the optimized kernels registered into their own table: 28 operators visible instead of 242. - Loading. A DLL records no search path. The Python entry points that load a DLL depending on executorch/lib register that directory first (portable_lib, kernels.quantized, codegen.tools, llm.custom_ops.op_tile_crop_aot), computed without resolving symlinks so an editable install, where those directories are links, registers its own lib too. A C++ consumer copies the DLLs beside its executable with `$<TARGET_RUNTIME_DLLS>`, which the C++ guide now shows, including in the quick start. - Debug. The DLLs use the C++ library of the configuration the wheel was built in, and a consumer in the other configuration would mix the two and corrupt memory. The package defines ET_PREBUILT_RELEASE_CRT, or ET_PREBUILT_DEBUG_CRT for a DEBUG=1 build, and the runtime headers refuse the mismatched configuration at compile time with a message naming the right one, for both single-config (-DCMAKE_BUILD_TYPE) and multi-config (--config) generators. naming the right one. (A /FAILIFMISMATCH link option would not do: link.exe ignores it on the command line and honours it only inside an object file.) - cpuinfo. The thread pool DLL exports pthreadpool, so the process has one pool, but keeps cpuinfo private. cpuinfo's feature checks (cpuinfo_has_x86_avx2 and the rest) are inline reads of cpuinfo_isa, which cpuinfo.h does not declare dllimport, so a DLL can only read its own copy; with cpuinfo exported, XNNPACK initialized the thread pool's copy and read its own, all zeros, and ran baseline kernels. Each DLL that uses cpuinfo now carries all of it. Measured on an x64 machine with AVX2: an XNNPACK-delegated 2 x Linear(1024) model at batch 256 went from 2.05-2.14 ms to 1.47-1.52 ms per run, with the same results. CUDA is not part of this. A Windows build with CUDA on keeps the static layout it had, since the Windows CUDA delegate has not been built as a DLL; a follow-up stacked on this ships it. Two bugs only the shared layout exposes on Windows are fixed here: - The process hung at exit about half the time. Windows terminates worker threads before running a DLL's static destructors, so destroying the global thread pool there waited on a lock a terminated worker could hold (`LdrShutdownProcess -> pthreadpool_destroy -> mtx_lock`). The pool is now deliberately not destroyed on Windows; Linux and macOS are unchanged. - `quantized_ops_aot_lib.dll` failed to load, which `executorch.kernels.quantized` swallows, so quantized export silently lost its out variants. The Windows wheel is built without the event tracer (unchanged), so the `etdump` component is not offered to C++ there rather than handing a consumer a profiler that records nothing. The delegates are unchanged too: the Windows wheel ships XNNPACK, as it did before this change. QNN, OpenVINO and TorchAO stay Linux (and for TorchAO, aarch64) wheel components, and Core ML and MLX macOS ones. QNN in particular builds on Windows (build-qnn-windows-x64 and -arm64 pass), but the wheel enables it only where pre_build_script.sh downloads the SDK and the pybind preset's Linux branch turns it on; bringing it to the Windows wheel means the SDK download on the Windows builder, its import library and anchor, and tests, which is a separate change. The pkg-config file the Linux and macOS wheels ship (#23169) is not generated on Windows either, since its link flags are GNU-style and a DLL records no search path for them; the test checks that it is absent. The pre-3.28 variables route links the import libraries and states C++20, which the Windows runtime headers need. ## Tests The two wheel suites now run on Windows, as they do on macOS since #21771, wired into test_windows.py. Their Windows forms ask the platform's own questions: - symbols: dumpbin /exports matched against MSVC decorated names, and for code a DLL takes from an archive, which the export list cannot name, the import direction: the owner must import `register_kernels` / `register_backend` from executorch.dll; - dependencies: dumpbin /dependents and /imports in place of readelf / otool; - loading: LoadLibrary per binary in its own process, which binds every import as ldd -r does, from the package and from a relocated copy; - paths: every recorded dependency is a bare DLL name; - platform tag: the PE machine field of every binary against win_amd64; - C++ consumers: built with --config Release, DLLs copied beside the executable, and the installed package removed from PATH so nothing resolves it for them; - package entry points: the extensions are imported the way a user reaches them (the pybindings through portable_lib), with nothing added to the DLL search path, and `import executorch.kernels.quantized` has to register the quantized out variants, so the entry points' own registration is what is tested; - a consumer in the wrong configuration has to be refused; an application using the thread pool has to exit five runs in a row; no DLL may import cpuinfo from another (the split state has no numeric symptom, only slower kernels); every program the suites build and every Python probe or export they start has a timeout, so a hang fails as a named timeout rather than the whole job timing out. Compiler, CMake and binary tool invocations are not bounded. ## Test plan On Windows 11 x64, VS 2022 BuildTools with ClangCL, CPython 3.12, from a wheel built with `python setup.py bdist_wheel` and installed into a clean environment: - `.ci/scripts/wheel/test_windows.py`: passes, including every check in test_shared_libraries.py and test_cpp_sdk.py (single owner of each component, _C contains no component, every DLL loads in place and relocated, custom op registers, platform tag; version request, 121 of 123 headers compile, documented example builds, runtime alone has no kernels, kernels / quantized / XNNPACK consumers match eager PyTorch, delegate fails without its component, pre-3.28 route on CMake 3.24, relocation, one shared registry) and MobileNetV3 through XNNPACK matching eager. - The exit hang: 8 of 20 and 10 of 20 runs of two consumers hung before the thread pool fix, 0 of 20 after. - Linux and macOS: the CMake changes are behind WIN32 / MSVC and the Python changes behind sys.platform == "win32"; relying on their wheel rows for confirmation.
Gasoonjia
added a commit
that referenced
this pull request
Oct 3, 2026
The Windows wheel shipped one fused Python extension and no linkable libraries, so a C++ application got nothing from it: the headers and the CMake package were there, but every component a consumer asked for resolved to nothing. Linux and macOS already ship the runtime, the kernels, the delegates and the thread pool as separate libraries that a C++ application links directly (#21610, #21771). This does the same for Windows. Before (Windows x64, CPython 3.12): extension/pybindings/_C.cp312-win_amd64.pyd 8.27 MB <- everything fused in here lib/ does not exist After: lib/executorch_kernels_optimized.dll 5.02 MB lib/executorch_backend_xnnpack.dll 2.47 MB lib/executorch.dll 0.38 MB lib/executorch_kernels_quantized.dll 0.22 MB lib/executorch_threadpool.dll 0.21 MB lib/executorch_etdump.dll 0.05 MB (loaded by _C; not offered to C++) lib/*.lib <- import libraries a consumer links extension/pybindings/_C.cp312-win_amd64.pyd 0.63 MB <- just the bindings now ## How The three mechanisms that made the shared layout Linux and macOS only now have Windows equivalents: what ELF / Mach-O Windows exporting the runtime API default visibility WINDOWS_EXPORT_ALL_SYMBOLS keeping a registration-only lib --no-as-needed / named dylib /INCLUDE of a per-DLL anchor finding sibling libraries $ORIGIN / @loader_path os.add_dll_directory, and a consumer copies the DLLs - Exports. The runtime has no dllexport annotations, which is why CMakeLists.txt refused a shared Windows build. Each shipped component DLL now exports all of its symbols, unless its target already exports through its own annotations (WINDOWS_EXPORT_ALL_SYMBOLS OFF), which it keeps. CMake builds that list from a target's own objects and cannot read an archive, so the runtime DLL takes its components as objects rather than through /WHOLEARCHIVE. The caller says which archives are part of a DLL's exports (the helper's EXPORTED option) instead of the helper inferring it from a property another call sets, so the result does not depend on call order. That uses $<COMPILE_ONLY>, new in CMake 3.27, so a shared Windows build asks for 3.27 with a clear message, the Windows build requirement is raised to match, and executorch_shared publishes the cxx_std_20 its headers need on Windows. The PAL source is built with its C functions strong there: the export list skips weak symbols, so the DLL exported none of the et_pal_* functions and clock.h's inline ticks_to_ns() failed to link against it. - Retention. A PE import survives only if some symbol from it is referenced, so each shipped DLL exports `executorch_anchor_<name>` and the imported targets carry `/INCLUDE` of it, the counterpart of the Linux `--no-as-needed`. The name follows the shipped file name without any configuration postfix, and is resolved at the end of configure and written out literally, so it survives `cmake --install` (a generator expression would be evaluated against the imported target, which has no OUTPUT_NAME) and is the same in every configuration of a multi-config build. - One registry. A consumer names the runtime's import library ahead of any archive, as on Linux, so nothing resolves the registry from a private static copy. Without this the optimized kernels registered into their own table: 28 operators visible instead of 242. - Loading. A DLL records no search path. The Python entry points that load a DLL depending on executorch/lib register that directory first (portable_lib, kernels.quantized, codegen.tools, llm.custom_ops.op_tile_crop_aot), computed without resolving symlinks so an editable install, where those directories are links, registers its own lib too. A C++ consumer copies the DLLs beside its executable with `$<TARGET_RUNTIME_DLLS>`, which the C++ guide now shows, including in the quick start. - Debug. The DLLs use the C++ library of the configuration the wheel was built in, and a consumer in the other configuration would mix the two and corrupt memory. The package defines ET_PREBUILT_RELEASE_CRT, or ET_PREBUILT_DEBUG_CRT for a DEBUG=1 build, and the runtime headers refuse the mismatched configuration at compile time with a message naming the right one, for both single-config (-DCMAKE_BUILD_TYPE) and multi-config (--config) generators. naming the right one. (A /FAILIFMISMATCH link option would not do: link.exe ignores it on the command line and honours it only inside an object file.) - cpuinfo. The thread pool DLL exports pthreadpool, so the process has one pool, but keeps cpuinfo private. cpuinfo's feature checks (cpuinfo_has_x86_avx2 and the rest) are inline reads of cpuinfo_isa, which cpuinfo.h does not declare dllimport, so a DLL can only read its own copy; with cpuinfo exported, XNNPACK initialized the thread pool's copy and read its own, all zeros, and ran baseline kernels. Each DLL that uses cpuinfo now carries all of it. Measured on an x64 machine with AVX2: an XNNPACK-delegated 2 x Linear(1024) model at batch 256 went from 2.05-2.14 ms to 1.47-1.52 ms per run, with the same results. CUDA is not part of this. A Windows build with CUDA on keeps the static layout it had, since the Windows CUDA delegate has not been built as a DLL; a follow-up stacked on this ships it. Two bugs only the shared layout exposes on Windows are fixed here: - The process hung at exit about half the time. Windows terminates worker threads before running a DLL's static destructors, so destroying the global thread pool there waited on a lock a terminated worker could hold (`LdrShutdownProcess -> pthreadpool_destroy -> mtx_lock`). The pool is now deliberately not destroyed on Windows; Linux and macOS are unchanged. - `quantized_ops_aot_lib.dll` failed to load, which `executorch.kernels.quantized` swallows, so quantized export silently lost its out variants. The Windows wheel is built without the event tracer (unchanged), so the `etdump` component is not offered to C++ there rather than handing a consumer a profiler that records nothing. The delegates are unchanged too: the Windows wheel ships XNNPACK, as it did before this change. QNN, OpenVINO and TorchAO stay Linux (and for TorchAO, aarch64) wheel components, and Core ML and MLX macOS ones. QNN in particular builds on Windows (build-qnn-windows-x64 and -arm64 pass), but the wheel enables it only where pre_build_script.sh downloads the SDK and the pybind preset's Linux branch turns it on; bringing it to the Windows wheel means the SDK download on the Windows builder, its import library and anchor, and tests, which is a separate change. The pkg-config file the Linux and macOS wheels ship (#23169) is not generated on Windows either, since its link flags are GNU-style and a DLL records no search path for them; the test checks that it is absent. The pre-3.28 variables route links the import libraries and states C++20, which the Windows runtime headers need. ## Tests The two wheel suites now run on Windows, as they do on macOS since #21771, wired into test_windows.py. Their Windows forms ask the platform's own questions: - symbols: dumpbin /exports matched against MSVC decorated names, and for code a DLL takes from an archive, which the export list cannot name, the import direction: the owner must import `register_kernels` / `register_backend` from executorch.dll; - dependencies: dumpbin /dependents and /imports in place of readelf / otool; - loading: LoadLibrary per binary in its own process, which binds every import as ldd -r does, from the package and from a relocated copy; - paths: every recorded dependency is a bare DLL name; - platform tag: the PE machine field of every binary against win_amd64; - C++ consumers: built with --config Release, DLLs copied beside the executable, and the installed package removed from PATH so nothing resolves it for them; - package entry points: the extensions are imported the way a user reaches them (the pybindings through portable_lib), with nothing added to the DLL search path, and `import executorch.kernels.quantized` has to register the quantized out variants, so the entry points' own registration is what is tested; - a consumer in the wrong configuration has to be refused; an application using the thread pool has to exit five runs in a row; no DLL may import cpuinfo from another (the split state has no numeric symptom, only slower kernels); every program the suites build and every Python probe or export they start has a timeout, so a hang fails as a named timeout rather than the whole job timing out. Compiler, CMake and binary tool invocations are not bounded. ## Test plan On Windows 11 x64, VS 2022 BuildTools with ClangCL, CPython 3.12, from a wheel built with `python setup.py bdist_wheel` and installed into a clean environment: - `.ci/scripts/wheel/test_windows.py`: passes, including every check in test_shared_libraries.py and test_cpp_sdk.py (single owner of each component, _C contains no component, every DLL loads in place and relocated, custom op registers, platform tag; version request, 121 of 123 headers compile, documented example builds, runtime alone has no kernels, kernels / quantized / XNNPACK consumers match eager PyTorch, delegate fails without its component, pre-3.28 route on CMake 3.24, relocation, one shared registry) and MobileNetV3 through XNNPACK matching eager. - The exit hang: 8 of 20 and 10 of 20 runs of two consumers hung before the thread pool fix, 0 of 20 after. - Linux and macOS: the CMake changes are behind WIN32 / MSVC and the Python changes behind sys.platform == "win32"; relying on their wheel rows for confirmation.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The wheel ships the runtime as shared libraries, with headers and a CMake package. Build systems that read pkg-config instead, such as Meson, look for
executorch.pc, and the wheel has none. So those projects need a full source build just to get that one file.This adds
lib/pkgconfig/executorch.pcto the Linux and macOS wheels:ET_EVENT_TRACER_ENABLEDwhen the libraries are built with the event tracer, andET_USE_THREADPOOLplus the thread pool library when the wheel ships the thread pool.LD_LIBRARY_PATH. On Linux the path is its own-Wl,-rpathargument, so Meson keeps it when it installs the program.runtimecomponent, it does not include the kernels. You name those on the link line.Usage on Linux:
On macOS, leave out the
--push-state,--no-as-neededand--pop-stateflags. The macOS linker keeps the kernels library anyway, and it rejects those flags.Windows is left out because the Windows wheel does not ship the runtime library.
The C++ usage page gets a pkg-config section with these commands and a Meson example. A new wheel smoke test builds a program with the flags pkg-config prints, runs a model, and compares the result with eager PyTorch. It also checks that the pkg-config file and the CMake package give the same compile definitions, so the two cannot drift apart. It gets pkg-config from the
pkgconfpackage on PyPI.Test plan: built the macOS arm64 wheel from this branch and installed it into a clean environment. The C++ SDK checks pass, including the new one. The new check fails when the
.pcfile is removed, and it fails when a compile definition is missing from it. With Meson on macOS and on Linux x86_64, the documented recipe builds, installs and runs a model, and the installed program still finds the libraries. A program that callsparallel_forruns on many threads when built from these flags. The Linux and CUDA wheels are covered by the wheel build jobs on this PR.