Skip to content

Ship a pkg-config file for the runtime in the Linux and macOS wheels - #23169

Merged
JakeStevens merged 4 commits into
pytorch:mainfrom
shoumikhin:wheel-pkg-config
Sep 26, 2026
Merged

JakeStevens merged 4 commits into
pytorch:mainfrom
shoumikhin:wheel-pkg-config

Conversation

@shoumikhin

@shoumikhin shoumikhin commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

The wheel ships the runtime as shared libraries, with headers and a CMake package. Build systems that read pkg-config instead, such as Meson, look for executorch.pc, and the wheel has none. So those projects need a full source build just to get that one file.

This adds lib/pkgconfig/executorch.pc to the Linux and macOS wheels:

  • All paths come from the file's own location, so it works wherever pip installed the wheel.
  • It gives the same compile definitions as the CMake package's runtime. That includes ET_EVENT_TRACER_ENABLED when the libraries are built with the event tracer, and ET_USE_THREADPOOL plus the thread pool library when the wheel ships the thread pool.
  • It adds a runtime search path, so a program finds the libraries without LD_LIBRARY_PATH. On Linux the path is its own -Wl,-rpath argument, so Meson keeps it when it installs the program.
  • Like the CMake runtime component, it does not include the kernels. You name those on the link line.

Usage on Linux:

export PKG_CONFIG_PATH="$(python -c 'import executorch, pathlib; print(pathlib.Path(executorch.__path__[0]) / "lib" / "pkgconfig")')"
c++ -std=c++17 main.cpp $(pkg-config --cflags --libs executorch) \
  -Wl,--push-state,--no-as-needed -lexecutorch_kernels_optimized -Wl,--pop-state -o app

On macOS, leave out the --push-state, --no-as-needed and --pop-state flags. The macOS linker keeps the kernels library anyway, and it rejects those flags.

Windows is left out because the Windows wheel does not ship the runtime library.

The C++ usage page gets a pkg-config section with these commands and a Meson example. A new wheel smoke test builds a program with the flags pkg-config prints, runs a model, and compares the result with eager PyTorch. It also checks that the pkg-config file and the CMake package give the same compile definitions, so the two cannot drift apart. It gets pkg-config from the pkgconf package on PyPI.

Test plan: built the macOS arm64 wheel from this branch and installed it into a clean environment. The C++ SDK checks pass, including the new one. The new check fails when the .pc file is removed, and it fails when a compile definition is missing from it. With Meson on macOS and on Linux x86_64, the documented recipe builds, installs and runs a model, and the installed program still finds the libraries. A program that calls parallel_for runs on many threads when built from these flags. The Linux and CUDA wheels are covered by the wheel build jobs on this PR.

The wheel ships the runtime as shared libraries with headers and a CMake
package, but build systems that read pkg-config, such as Meson, cannot use it.
They look for executorch.pc and the wheel has none, so today they need a full
source build and install just to get that one file.

This adds lib/pkgconfig/executorch.pc to the wheel. It works out every path
from its own location, so the wheel can be installed anywhere. It passes the
same compile definitions as the CMake package, including the event tracer
switch when the libraries are built with it, and it adds a runtime search path
so a program finds the libraries without LD_LIBRARY_PATH. Like the CMake
runtime component, it covers the runtime only, and the kernel libraries are
named on the link line.

Windows is left out because the Windows wheel does not ship the runtime
library.
Copilot AI lite review requested due to automatic review settings September 25, 2026 21:48
@shoumikhin shoumikhin added the release notes: build Changes related to build, including dependency upgrades, build flags, optimizations, etc. label Sep 25, 2026
@pytorch-bot

pytorch-bot Bot commented Sep 25, 2026 •

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/23169

Note: Links to docs will display an error until the docs builds have been completed.

✅ No Failures

As of commit fd6587f with merge base 4199643 (image):
💚 Looks good so far! There are no failures yet. 💚

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-cla meta-cla Bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Sep 25, 2026

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

Copilot AI review requested due to automatic review settings September 25, 2026 23:59

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

Copilot AI review requested due to automatic review settings September 26, 2026 02:27

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

The pkg-config file left out the thread pool. The CMake package adds
ET_USE_THREADPOOL and the thread pool library to the runtime whenever the
wheel ships the thread pool. Without them, code built from the pkg-config
flags gets the serial fallback of parallel_for, with no error. The file now
adds both under the same condition.

On Linux the runtime search path was joined to --enable-new-dtags in one
linker argument. Meson only keeps a dependency's search path when the
argument starts with -Wl,-rpath, so meson install removed it and the
installed program could not find the libraries. They are now two arguments.

The smoke test now also checks that the compile definitions pkg-config
prints match the ones the CMake package gives, so the two cannot drift.

The docs now show separate Linux and macOS commands, scope --no-as-needed
with push-state and pop-state, add to PKG_CONFIG_PATH instead of replacing
it, ask for an absolute path, and give a Meson example that was built,
installed and run.
Copilot AI review requested due to automatic review settings September 26, 2026 03:15

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@JakeStevens
JakeStevens merged commit b2896ae into pytorch:main Sep 26, 2026
388 checks passed
Gasoonjia added a commit that referenced this pull request Sep 29, 2026
The Windows wheel shipped one fused Python extension and no linkable libraries, so a C++
application got nothing from it: the headers and the CMake package were there, but every
component a consumer asked for resolved to nothing. Linux and macOS already ship the runtime,
the kernels, the delegates and the thread pool as separate libraries that a C++ application
links directly (#21610, #21771). This does the same for Windows.

Before (Windows x64, CPython 3.12):

    extension/pybindings/_C.cp312-win_amd64.pyd   8.27 MB   <- everything fused in here
    lib/                                          does not exist

After:

    lib/executorch_kernels_optimized.dll          5.02 MB
    lib/executorch_backend_xnnpack.dll            2.47 MB
    lib/executorch.dll                            0.38 MB
    lib/executorch_kernels_quantized.dll          0.22 MB
    lib/executorch_threadpool.dll                 0.21 MB
    lib/executorch_etdump.dll                     0.05 MB   (loaded by _C; not offered to C++)
    lib/*.lib                                               <- import libraries a consumer links
    extension/pybindings/_C.cp312-win_amd64.pyd   0.63 MB   <- just the bindings now

## How

The three mechanisms that made the shared layout Linux and macOS only now have Windows
equivalents:

    what                              ELF / Mach-O                   Windows
    exporting the runtime API         default visibility             WINDOWS_EXPORT_ALL_SYMBOLS
    keeping a registration-only lib   --no-as-needed / named dylib   /INCLUDE of a per-DLL anchor
    finding sibling libraries         $ORIGIN / @loader_path         os.add_dll_directory, and a
                                                                     consumer copies the DLLs

- Exports. The runtime has no dllexport annotations, which is why CMakeLists.txt refused a
  shared Windows build. Each shipped component DLL now exports all of its symbols, unless
  its target already exports through its own annotations (WINDOWS_EXPORT_ALL_SYMBOLS OFF),
  which it keeps. CMake builds that list from a target's own objects and cannot read an
  archive, so the runtime DLL takes its components as objects rather than through
  /WHOLEARCHIVE. That uses $<COMPILE_ONLY>, new in CMake 3.27, so a shared Windows build
  asks for 3.27 with a clear message, and the Windows build requirement is raised to match.
- Retention. A PE import survives only if some symbol from it is referenced, so each
  shipped DLL exports `executorch_anchor_<name>` and the imported targets carry
  `/INCLUDE` of it, the counterpart of the Linux `--no-as-needed`.
- One registry. A consumer names the runtime's import library ahead of any archive, as on
  Linux, so nothing resolves the registry from a private static copy. Without this the
  optimized kernels registered into their own table: 28 operators visible instead of 242.
- Loading. A DLL records no search path. The Python entry points that load a DLL depending
  on executorch/lib register that directory first (portable_lib, kernels.quantized,
  codegen.tools, llm.custom_ops.op_tile_crop_aot), computed without resolving symlinks so
  an editable install, where those directories are links, registers its own lib too. A C++
  consumer copies the DLLs beside its executable with `$<TARGET_RUNTIME_DLLS>`, which the
  C++ guide now shows, including in the quick start.
- Debug. The DLLs use the release C++ library, and a Debug consumer mixing in the debug one
  would corrupt memory. The package defines ET_PREBUILT_RELEASE_CRT for its consumers, and
  the runtime headers refuse a Debug build with it at compile time, with a message asking
  for Release. (A /FAILIFMISMATCH link option would not do: link.exe ignores it on the
  command line and honours it only inside an object file.)

CUDA is not part of this. A Windows build with CUDA on keeps the static layout it had,
since the Windows CUDA delegate has not been built as a DLL; a follow-up stacked on this
ships it.

Two bugs only the shared layout exposes on Windows are fixed here:

- The process hung at exit about half the time. Windows terminates worker threads before
  running a DLL's static destructors, so destroying the global thread pool there waited on a
  lock a terminated worker could hold (`LdrShutdownProcess -> pthreadpool_destroy ->
  mtx_lock`). The pool is now deliberately not destroyed on Windows; Linux and macOS are
  unchanged.
- `quantized_ops_aot_lib.dll` failed to load, which `executorch.kernels.quantized` swallows,
  so quantized export silently lost its out variants.

The Windows wheel is built without the event tracer (unchanged), so the `etdump` component
is not offered to C++ there rather than handing a consumer a profiler that records nothing.
The pkg-config file the Linux and macOS wheels ship (#23169) is not generated on Windows
either, since its link flags are GNU-style and a DLL records no search path for them; the
test checks that it is absent.
The pre-3.28 variables route links the import libraries and states C++20, which the Windows
runtime headers need.

## Tests

The two wheel suites now run on Windows, as they do on macOS since #21771, wired into
test_windows.py. Their Windows forms ask the platform's own questions:

- symbols: dumpbin /exports matched against MSVC decorated names, and for code a DLL takes
  from an archive, which the export list cannot name, the import direction: the owner must
  import `register_kernels` / `register_backend` from executorch.dll;
- dependencies: dumpbin /dependents and /imports in place of readelf / otool;
- loading: LoadLibrary per binary in its own process, which binds every import as ldd -r
  does, from the package and from a relocated copy;
- paths: every recorded dependency is a bare DLL name;
- platform tag: the PE machine field of every binary against win_amd64;
- C++ consumers: built with --config Release, DLLs copied beside the executable, and the
  installed package removed from PATH so nothing resolves it for them;
- package entry points: the extensions are imported the way a user reaches them (the
  pybindings through portable_lib), with nothing added to the DLL search path, and
  `import executorch.kernels.quantized` has to register the quantized out variants, so the
  entry points' own registration is what is tested;
- a Debug consumer has to be refused; an application using the thread pool has to exit
  five runs in a row; every program and probe the suites run has a timeout, so a hang fails
  as a named timeout rather than the whole job timing out.

## Test plan

On Windows 11 x64, VS 2022 BuildTools with ClangCL, CPython 3.12, from a wheel built with
`python setup.py bdist_wheel` and installed into a clean environment:

- `.ci/scripts/wheel/test_windows.py`: passes, including every check in
  test_shared_libraries.py and test_cpp_sdk.py (single owner of each component, _C contains
  no component, every DLL loads in place and relocated, custom op registers, platform tag;
  version request, 121 of 123 headers compile, documented example builds, runtime alone has
  no kernels, kernels / quantized / XNNPACK consumers match eager PyTorch, delegate fails
  without its component, pre-3.28 route on CMake 3.24, relocation, one shared registry) and
  MobileNetV3 through XNNPACK matching eager.
- The exit hang: 8 of 20 and 10 of 20 runs of two consumers hung before the thread pool fix,
  0 of 20 after.
- Linux and macOS: the CMake changes are behind WIN32 / MSVC and the Python changes behind
  sys.platform == "win32"; relying on their wheel rows for confirmation.
Gasoonjia added a commit that referenced this pull request Sep 30, 2026
The Windows wheel shipped one fused Python extension and no linkable libraries, so a C++
application got nothing from it: the headers and the CMake package were there, but every
component a consumer asked for resolved to nothing. Linux and macOS already ship the runtime,
the kernels, the delegates and the thread pool as separate libraries that a C++ application
links directly (#21610, #21771). This does the same for Windows.

Before (Windows x64, CPython 3.12):

    extension/pybindings/_C.cp312-win_amd64.pyd   8.27 MB   <- everything fused in here
    lib/                                          does not exist

After:

    lib/executorch_kernels_optimized.dll          5.02 MB
    lib/executorch_backend_xnnpack.dll            2.47 MB
    lib/executorch.dll                            0.38 MB
    lib/executorch_kernels_quantized.dll          0.22 MB
    lib/executorch_threadpool.dll                 0.21 MB
    lib/executorch_etdump.dll                     0.05 MB   (loaded by _C; not offered to C++)
    lib/*.lib                                               <- import libraries a consumer links
    extension/pybindings/_C.cp312-win_amd64.pyd   0.63 MB   <- just the bindings now

## How

The three mechanisms that made the shared layout Linux and macOS only now have Windows
equivalents:

    what                              ELF / Mach-O                   Windows
    exporting the runtime API         default visibility             WINDOWS_EXPORT_ALL_SYMBOLS
    keeping a registration-only lib   --no-as-needed / named dylib   /INCLUDE of a per-DLL anchor
    finding sibling libraries         $ORIGIN / @loader_path         os.add_dll_directory, and a
                                                                     consumer copies the DLLs

- Exports. The runtime has no dllexport annotations, which is why CMakeLists.txt refused a
  shared Windows build. Each shipped component DLL now exports all of its symbols, unless
  its target already exports through its own annotations (WINDOWS_EXPORT_ALL_SYMBOLS OFF),
  which it keeps. CMake builds that list from a target's own objects and cannot read an
  archive, so the runtime DLL takes its components as objects rather than through
  /WHOLEARCHIVE. The caller says which archives are part of a DLL's exports (the helper's
  EXPORTED option) instead of the helper inferring it from a property another call sets,
  so the result does not depend on call order. That uses $<COMPILE_ONLY>, new in CMake
  3.27, so a shared Windows build asks for 3.27 with a clear message, the Windows build
  requirement is raised to match, and executorch_shared publishes the cxx_std_20 its
  headers need on Windows.
- Retention. A PE import survives only if some symbol from it is referenced, so each
  shipped DLL exports `executorch_anchor_<name>` and the imported targets carry
  `/INCLUDE` of it, the counterpart of the Linux `--no-as-needed`. The anchor source is
  generated per configuration, so a multi-config build with CMAKE_DEBUG_POSTFIX works.
- One registry. A consumer names the runtime's import library ahead of any archive, as on
  Linux, so nothing resolves the registry from a private static copy. Without this the
  optimized kernels registered into their own table: 28 operators visible instead of 242.
- Loading. A DLL records no search path. The Python entry points that load a DLL depending
  on executorch/lib register that directory first (portable_lib, kernels.quantized,
  codegen.tools, llm.custom_ops.op_tile_crop_aot), computed without resolving symlinks so
  an editable install, where those directories are links, registers its own lib too. A C++
  consumer copies the DLLs beside its executable with `$<TARGET_RUNTIME_DLLS>`, which the
  C++ guide now shows, including in the quick start.
- Debug. The DLLs use the C++ library of the configuration the wheel was built in, and a
  consumer in the other configuration would mix the two and corrupt memory. The package
  defines ET_PREBUILT_RELEASE_CRT, or ET_PREBUILT_DEBUG_CRT for a DEBUG=1 build, and the
  runtime headers refuse the mismatched configuration at compile time with a message
  naming the right one. (A /FAILIFMISMATCH link option would not do: link.exe ignores it
  on the command line and honours it only inside an object file.)
- cpuinfo. The thread pool DLL exports pthreadpool, so the process has one pool, but keeps
  cpuinfo private. cpuinfo's feature checks (cpuinfo_has_x86_avx2 and the rest) are
  inline reads of cpuinfo_isa, which cpuinfo.h does not declare dllimport, so a DLL can
  only read its own copy; with cpuinfo exported, XNNPACK initialized the thread pool's copy
  and read its own, all zeros, and ran baseline kernels. Each DLL that uses cpuinfo now
  carries all of it. Measured on an x64 machine with AVX2: an XNNPACK-delegated
  2 x Linear(1024) model at batch 256 went from 2.05-2.14 ms to 1.47-1.52 ms per run, with
  the same results.

CUDA is not part of this. A Windows build with CUDA on keeps the static layout it had,
since the Windows CUDA delegate has not been built as a DLL; a follow-up stacked on this
ships it.

Two bugs only the shared layout exposes on Windows are fixed here:

- The process hung at exit about half the time. Windows terminates worker threads before
  running a DLL's static destructors, so destroying the global thread pool there waited on a
  lock a terminated worker could hold (`LdrShutdownProcess -> pthreadpool_destroy ->
  mtx_lock`). The pool is now deliberately not destroyed on Windows; Linux and macOS are
  unchanged.
- `quantized_ops_aot_lib.dll` failed to load, which `executorch.kernels.quantized` swallows,
  so quantized export silently lost its out variants.

The Windows wheel is built without the event tracer (unchanged), so the `etdump` component
is not offered to C++ there rather than handing a consumer a profiler that records nothing.
The pkg-config file the Linux and macOS wheels ship (#23169) is not generated on Windows
either, since its link flags are GNU-style and a DLL records no search path for them; the
test checks that it is absent.
The pre-3.28 variables route links the import libraries and states C++20, which the Windows
runtime headers need.

## Tests

The two wheel suites now run on Windows, as they do on macOS since #21771, wired into
test_windows.py. Their Windows forms ask the platform's own questions:

- symbols: dumpbin /exports matched against MSVC decorated names, and for code a DLL takes
  from an archive, which the export list cannot name, the import direction: the owner must
  import `register_kernels` / `register_backend` from executorch.dll;
- dependencies: dumpbin /dependents and /imports in place of readelf / otool;
- loading: LoadLibrary per binary in its own process, which binds every import as ldd -r
  does, from the package and from a relocated copy;
- paths: every recorded dependency is a bare DLL name;
- platform tag: the PE machine field of every binary against win_amd64;
- C++ consumers: built with --config Release, DLLs copied beside the executable, and the
  installed package removed from PATH so nothing resolves it for them;
- package entry points: the extensions are imported the way a user reaches them (the
  pybindings through portable_lib), with nothing added to the DLL search path, and
  `import executorch.kernels.quantized` has to register the quantized out variants, so the
  entry points' own registration is what is tested;
- a consumer in the wrong configuration has to be refused; an application using the
  thread pool has to exit five runs in a row; no DLL may import cpuinfo from another
  (the split state has no numeric symptom, only slower kernels); every program the suites
  build and every Python probe or export they start has a timeout, so a hang fails as a
  named timeout rather than the whole job timing out. Compiler, CMake and binary tool
  invocations are not bounded.

## Test plan

On Windows 11 x64, VS 2022 BuildTools with ClangCL, CPython 3.12, from a wheel built with
`python setup.py bdist_wheel` and installed into a clean environment:

- `.ci/scripts/wheel/test_windows.py`: passes, including every check in
  test_shared_libraries.py and test_cpp_sdk.py (single owner of each component, _C contains
  no component, every DLL loads in place and relocated, custom op registers, platform tag;
  version request, 121 of 123 headers compile, documented example builds, runtime alone has
  no kernels, kernels / quantized / XNNPACK consumers match eager PyTorch, delegate fails
  without its component, pre-3.28 route on CMake 3.24, relocation, one shared registry) and
  MobileNetV3 through XNNPACK matching eager.
- The exit hang: 8 of 20 and 10 of 20 runs of two consumers hung before the thread pool fix,
  0 of 20 after.
- Linux and macOS: the CMake changes are behind WIN32 / MSVC and the Python changes behind
  sys.platform == "win32"; relying on their wheel rows for confirmation.
Gasoonjia added a commit that referenced this pull request Oct 2, 2026
The Windows wheel shipped one fused Python extension and no linkable libraries, so a C++
application got nothing from it: the headers and the CMake package were there, but every
component a consumer asked for resolved to nothing. Linux and macOS already ship the runtime,
the kernels, the delegates and the thread pool as separate libraries that a C++ application
links directly (#21610, #21771). This does the same for Windows.

Before (Windows x64, CPython 3.12):

    extension/pybindings/_C.cp312-win_amd64.pyd   8.27 MB   <- everything fused in here
    lib/                                          does not exist

After:

    lib/executorch_kernels_optimized.dll          5.02 MB
    lib/executorch_backend_xnnpack.dll            2.47 MB
    lib/executorch.dll                            0.38 MB
    lib/executorch_kernels_quantized.dll          0.22 MB
    lib/executorch_threadpool.dll                 0.21 MB
    lib/executorch_etdump.dll                     0.05 MB   (loaded by _C; not offered to C++)
    lib/*.lib                                               <- import libraries a consumer links
    extension/pybindings/_C.cp312-win_amd64.pyd   0.63 MB   <- just the bindings now

## How

The three mechanisms that made the shared layout Linux and macOS only now have Windows
equivalents:

    what                              ELF / Mach-O                   Windows
    exporting the runtime API         default visibility             WINDOWS_EXPORT_ALL_SYMBOLS
    keeping a registration-only lib   --no-as-needed / named dylib   /INCLUDE of a per-DLL anchor
    finding sibling libraries         $ORIGIN / @loader_path         os.add_dll_directory, and a
                                                                     consumer copies the DLLs

- Exports. The runtime has no dllexport annotations, which is why CMakeLists.txt refused a
  shared Windows build. Each shipped component DLL now exports all of its symbols, unless
  its target already exports through its own annotations (WINDOWS_EXPORT_ALL_SYMBOLS OFF),
  which it keeps. CMake builds that list from a target's own objects and cannot read an
  archive, so the runtime DLL takes its components as objects rather than through
  /WHOLEARCHIVE. The caller says which archives are part of a DLL's exports (the helper's
  EXPORTED option) instead of the helper inferring it from a property another call sets,
  so the result does not depend on call order. That uses $<COMPILE_ONLY>, new in CMake
  3.27, so a shared Windows build asks for 3.27 with a clear message, the Windows build
  requirement is raised to match, and executorch_shared publishes the cxx_std_20 its
  headers need on Windows. The PAL source is built with its C functions strong there:
  the export list skips weak symbols, so the DLL exported none of the et_pal_* functions
  and clock.h's inline ticks_to_ns() failed to link against it.
- Retention. A PE import survives only if some symbol from it is referenced, so each
  shipped DLL exports `executorch_anchor_<name>` and the imported targets carry
  `/INCLUDE` of it, the counterpart of the Linux `--no-as-needed`. The name follows the
  shipped file name without any configuration postfix, and is resolved at the end of
  configure and written out literally, so it survives `cmake --install` (a generator
  expression would be evaluated against the imported target, which has no OUTPUT_NAME)
  and is the same in every configuration of a multi-config build.
- One registry. A consumer names the runtime's import library ahead of any archive, as on
  Linux, so nothing resolves the registry from a private static copy. Without this the
  optimized kernels registered into their own table: 28 operators visible instead of 242.
- Loading. A DLL records no search path. The Python entry points that load a DLL depending
  on executorch/lib register that directory first (portable_lib, kernels.quantized,
  codegen.tools, llm.custom_ops.op_tile_crop_aot), computed without resolving symlinks so
  an editable install, where those directories are links, registers its own lib too. A C++
  consumer copies the DLLs beside its executable with `$<TARGET_RUNTIME_DLLS>`, which the
  C++ guide now shows, including in the quick start.
- Debug. The DLLs use the C++ library of the configuration the wheel was built in, and a
  consumer in the other configuration would mix the two and corrupt memory. The package
  defines ET_PREBUILT_RELEASE_CRT, or ET_PREBUILT_DEBUG_CRT for a DEBUG=1 build, and the
  runtime headers refuse the mismatched configuration at compile time with a message
  naming the right one, for both single-config (-DCMAKE_BUILD_TYPE) and multi-config
  (--config) generators.
  naming the right one. (A /FAILIFMISMATCH link option would not do: link.exe ignores it
  on the command line and honours it only inside an object file.)
- cpuinfo. The thread pool DLL exports pthreadpool, so the process has one pool, but keeps
  cpuinfo private. cpuinfo's feature checks (cpuinfo_has_x86_avx2 and the rest) are
  inline reads of cpuinfo_isa, which cpuinfo.h does not declare dllimport, so a DLL can
  only read its own copy; with cpuinfo exported, XNNPACK initialized the thread pool's copy
  and read its own, all zeros, and ran baseline kernels. Each DLL that uses cpuinfo now
  carries all of it. Measured on an x64 machine with AVX2: an XNNPACK-delegated
  2 x Linear(1024) model at batch 256 went from 2.05-2.14 ms to 1.47-1.52 ms per run, with
  the same results.

CUDA is not part of this. A Windows build with CUDA on keeps the static layout it had,
since the Windows CUDA delegate has not been built as a DLL; a follow-up stacked on this
ships it.

Two bugs only the shared layout exposes on Windows are fixed here:

- The process hung at exit about half the time. Windows terminates worker threads before
  running a DLL's static destructors, so destroying the global thread pool there waited on a
  lock a terminated worker could hold (`LdrShutdownProcess -> pthreadpool_destroy ->
  mtx_lock`). The pool is now deliberately not destroyed on Windows; Linux and macOS are
  unchanged.
- `quantized_ops_aot_lib.dll` failed to load, which `executorch.kernels.quantized` swallows,
  so quantized export silently lost its out variants.

The Windows wheel is built without the event tracer (unchanged), so the `etdump` component
is not offered to C++ there rather than handing a consumer a profiler that records nothing.

The delegates are unchanged too: the Windows wheel ships XNNPACK, as it did before this
change. QNN, OpenVINO and TorchAO stay Linux (and for TorchAO, aarch64) wheel components, and
Core ML and MLX macOS ones. QNN in particular builds on Windows (build-qnn-windows-x64 and
-arm64 pass), but the wheel enables it only where pre_build_script.sh downloads the SDK and
the pybind preset's Linux branch turns it on; bringing it to the Windows wheel means the SDK
download on the Windows builder, its import library and anchor, and tests, which is a
separate change.
The pkg-config file the Linux and macOS wheels ship (#23169) is not generated on Windows
either, since its link flags are GNU-style and a DLL records no search path for them; the
test checks that it is absent.
The pre-3.28 variables route links the import libraries and states C++20, which the Windows
runtime headers need.

## Tests

The two wheel suites now run on Windows, as they do on macOS since #21771, wired into
test_windows.py. Their Windows forms ask the platform's own questions:

- symbols: dumpbin /exports matched against MSVC decorated names, and for code a DLL takes
  from an archive, which the export list cannot name, the import direction: the owner must
  import `register_kernels` / `register_backend` from executorch.dll;
- dependencies: dumpbin /dependents and /imports in place of readelf / otool;
- loading: LoadLibrary per binary in its own process, which binds every import as ldd -r
  does, from the package and from a relocated copy;
- paths: every recorded dependency is a bare DLL name;
- platform tag: the PE machine field of every binary against win_amd64;
- C++ consumers: built with --config Release, DLLs copied beside the executable, and the
  installed package removed from PATH so nothing resolves it for them;
- package entry points: the extensions are imported the way a user reaches them (the
  pybindings through portable_lib), with nothing added to the DLL search path, and
  `import executorch.kernels.quantized` has to register the quantized out variants, so the
  entry points' own registration is what is tested;
- a consumer in the wrong configuration has to be refused; an application using the
  thread pool has to exit five runs in a row; no DLL may import cpuinfo from another
  (the split state has no numeric symptom, only slower kernels); every program the suites
  build and every Python probe or export they start has a timeout, so a hang fails as a
  named timeout rather than the whole job timing out. Compiler, CMake and binary tool
  invocations are not bounded.

## Test plan

On Windows 11 x64, VS 2022 BuildTools with ClangCL, CPython 3.12, from a wheel built with
`python setup.py bdist_wheel` and installed into a clean environment:

- `.ci/scripts/wheel/test_windows.py`: passes, including every check in
  test_shared_libraries.py and test_cpp_sdk.py (single owner of each component, _C contains
  no component, every DLL loads in place and relocated, custom op registers, platform tag;
  version request, 121 of 123 headers compile, documented example builds, runtime alone has
  no kernels, kernels / quantized / XNNPACK consumers match eager PyTorch, delegate fails
  without its component, pre-3.28 route on CMake 3.24, relocation, one shared registry) and
  MobileNetV3 through XNNPACK matching eager.
- The exit hang: 8 of 20 and 10 of 20 runs of two consumers hung before the thread pool fix,
  0 of 20 after.
- Linux and macOS: the CMake changes are behind WIN32 / MSVC and the Python changes behind
  sys.platform == "win32"; relying on their wheel rows for confirmation.
Gasoonjia added a commit that referenced this pull request Oct 3, 2026
The Windows wheel shipped one fused Python extension and no linkable libraries, so a C++
application got nothing from it: the headers and the CMake package were there, but every
component a consumer asked for resolved to nothing. Linux and macOS already ship the runtime,
the kernels, the delegates and the thread pool as separate libraries that a C++ application
links directly (#21610, #21771). This does the same for Windows.

Before (Windows x64, CPython 3.12):

    extension/pybindings/_C.cp312-win_amd64.pyd   8.27 MB   <- everything fused in here
    lib/                                          does not exist

After:

    lib/executorch_kernels_optimized.dll          5.02 MB
    lib/executorch_backend_xnnpack.dll            2.47 MB
    lib/executorch.dll                            0.38 MB
    lib/executorch_kernels_quantized.dll          0.22 MB
    lib/executorch_threadpool.dll                 0.21 MB
    lib/executorch_etdump.dll                     0.05 MB   (loaded by _C; not offered to C++)
    lib/*.lib                                               <- import libraries a consumer links
    extension/pybindings/_C.cp312-win_amd64.pyd   0.63 MB   <- just the bindings now

## How

The three mechanisms that made the shared layout Linux and macOS only now have Windows
equivalents:

    what                              ELF / Mach-O                   Windows
    exporting the runtime API         default visibility             WINDOWS_EXPORT_ALL_SYMBOLS
    keeping a registration-only lib   --no-as-needed / named dylib   /INCLUDE of a per-DLL anchor
    finding sibling libraries         $ORIGIN / @loader_path         os.add_dll_directory, and a
                                                                     consumer copies the DLLs

- Exports. The runtime has no dllexport annotations, which is why CMakeLists.txt refused a
  shared Windows build. Each shipped component DLL now exports all of its symbols, unless
  its target already exports through its own annotations (WINDOWS_EXPORT_ALL_SYMBOLS OFF),
  which it keeps. CMake builds that list from a target's own objects and cannot read an
  archive, so the runtime DLL takes its components as objects rather than through
  /WHOLEARCHIVE. The caller says which archives are part of a DLL's exports (the helper's
  EXPORTED option) instead of the helper inferring it from a property another call sets,
  so the result does not depend on call order. That uses $<COMPILE_ONLY>, new in CMake
  3.27, so a shared Windows build asks for 3.27 with a clear message, the Windows build
  requirement is raised to match, and executorch_shared publishes the cxx_std_20 its
  headers need on Windows. The PAL source is built with its C functions strong there:
  the export list skips weak symbols, so the DLL exported none of the et_pal_* functions
  and clock.h's inline ticks_to_ns() failed to link against it.
- Retention. A PE import survives only if some symbol from it is referenced, so each
  shipped DLL exports `executorch_anchor_<name>` and the imported targets carry
  `/INCLUDE` of it, the counterpart of the Linux `--no-as-needed`. The name follows the
  shipped file name without any configuration postfix, and is resolved at the end of
  configure and written out literally, so it survives `cmake --install` (a generator
  expression would be evaluated against the imported target, which has no OUTPUT_NAME)
  and is the same in every configuration of a multi-config build.
- One registry. A consumer names the runtime's import library ahead of any archive, as on
  Linux, so nothing resolves the registry from a private static copy. Without this the
  optimized kernels registered into their own table: 28 operators visible instead of 242.
- Loading. A DLL records no search path. The Python entry points that load a DLL depending
  on executorch/lib register that directory first (portable_lib, kernels.quantized,
  codegen.tools, llm.custom_ops.op_tile_crop_aot), computed without resolving symlinks so
  an editable install, where those directories are links, registers its own lib too. A C++
  consumer copies the DLLs beside its executable with `$<TARGET_RUNTIME_DLLS>`, which the
  C++ guide now shows, including in the quick start.
- Debug. The DLLs use the C++ library of the configuration the wheel was built in, and a
  consumer in the other configuration would mix the two and corrupt memory. The package
  defines ET_PREBUILT_RELEASE_CRT, or ET_PREBUILT_DEBUG_CRT for a DEBUG=1 build, and the
  runtime headers refuse the mismatched configuration at compile time with a message
  naming the right one, for both single-config (-DCMAKE_BUILD_TYPE) and multi-config
  (--config) generators.
  naming the right one. (A /FAILIFMISMATCH link option would not do: link.exe ignores it
  on the command line and honours it only inside an object file.)
- cpuinfo. The thread pool DLL exports pthreadpool, so the process has one pool, but keeps
  cpuinfo private. cpuinfo's feature checks (cpuinfo_has_x86_avx2 and the rest) are
  inline reads of cpuinfo_isa, which cpuinfo.h does not declare dllimport, so a DLL can
  only read its own copy; with cpuinfo exported, XNNPACK initialized the thread pool's copy
  and read its own, all zeros, and ran baseline kernels. Each DLL that uses cpuinfo now
  carries all of it. Measured on an x64 machine with AVX2: an XNNPACK-delegated
  2 x Linear(1024) model at batch 256 went from 2.05-2.14 ms to 1.47-1.52 ms per run, with
  the same results.

CUDA is not part of this. A Windows build with CUDA on keeps the static layout it had,
since the Windows CUDA delegate has not been built as a DLL; a follow-up stacked on this
ships it.

Two bugs only the shared layout exposes on Windows are fixed here:

- The process hung at exit about half the time. Windows terminates worker threads before
  running a DLL's static destructors, so destroying the global thread pool there waited on a
  lock a terminated worker could hold (`LdrShutdownProcess -> pthreadpool_destroy ->
  mtx_lock`). The pool is now deliberately not destroyed on Windows; Linux and macOS are
  unchanged.
- `quantized_ops_aot_lib.dll` failed to load, which `executorch.kernels.quantized` swallows,
  so quantized export silently lost its out variants.

The Windows wheel is built without the event tracer (unchanged), so the `etdump` component
is not offered to C++ there rather than handing a consumer a profiler that records nothing.

The delegates are unchanged too: the Windows wheel ships XNNPACK, as it did before this
change. QNN, OpenVINO and TorchAO stay Linux (and for TorchAO, aarch64) wheel components, and
Core ML and MLX macOS ones. QNN in particular builds on Windows (build-qnn-windows-x64 and
-arm64 pass), but the wheel enables it only where pre_build_script.sh downloads the SDK and
the pybind preset's Linux branch turns it on; bringing it to the Windows wheel means the SDK
download on the Windows builder, its import library and anchor, and tests, which is a
separate change.
The pkg-config file the Linux and macOS wheels ship (#23169) is not generated on Windows
either, since its link flags are GNU-style and a DLL records no search path for them; the
test checks that it is absent.
The pre-3.28 variables route links the import libraries and states C++20, which the Windows
runtime headers need.

## Tests

The two wheel suites now run on Windows, as they do on macOS since #21771, wired into
test_windows.py. Their Windows forms ask the platform's own questions:

- symbols: dumpbin /exports matched against MSVC decorated names, and for code a DLL takes
  from an archive, which the export list cannot name, the import direction: the owner must
  import `register_kernels` / `register_backend` from executorch.dll;
- dependencies: dumpbin /dependents and /imports in place of readelf / otool;
- loading: LoadLibrary per binary in its own process, which binds every import as ldd -r
  does, from the package and from a relocated copy;
- paths: every recorded dependency is a bare DLL name;
- platform tag: the PE machine field of every binary against win_amd64;
- C++ consumers: built with --config Release, DLLs copied beside the executable, and the
  installed package removed from PATH so nothing resolves it for them;
- package entry points: the extensions are imported the way a user reaches them (the
  pybindings through portable_lib), with nothing added to the DLL search path, and
  `import executorch.kernels.quantized` has to register the quantized out variants, so the
  entry points' own registration is what is tested;
- a consumer in the wrong configuration has to be refused; an application using the
  thread pool has to exit five runs in a row; no DLL may import cpuinfo from another
  (the split state has no numeric symptom, only slower kernels); every program the suites
  build and every Python probe or export they start has a timeout, so a hang fails as a
  named timeout rather than the whole job timing out. Compiler, CMake and binary tool
  invocations are not bounded.

## Test plan

On Windows 11 x64, VS 2022 BuildTools with ClangCL, CPython 3.12, from a wheel built with
`python setup.py bdist_wheel` and installed into a clean environment:

- `.ci/scripts/wheel/test_windows.py`: passes, including every check in
  test_shared_libraries.py and test_cpp_sdk.py (single owner of each component, _C contains
  no component, every DLL loads in place and relocated, custom op registers, platform tag;
  version request, 121 of 123 headers compile, documented example builds, runtime alone has
  no kernels, kernels / quantized / XNNPACK consumers match eager PyTorch, delegate fails
  without its component, pre-3.28 route on CMake 3.24, relocation, one shared registry) and
  MobileNetV3 through XNNPACK matching eager.
- The exit hang: 8 of 20 and 10 of 20 runs of two consumers hung before the thread pool fix,
  0 of 20 after.
- Linux and macOS: the CMake changes are behind WIN32 / MSVC and the Python changes behind
  sys.platform == "win32"; relying on their wheel rows for confirmation.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. release notes: build Changes related to build, including dependency upgrades, build flags, optimizations, etc.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants