Skip to content

Document Numba/OMP_NUM_THREADS oversubscription alongside FFT threading - #29

Merged
TomaSusi merged 2 commits into
mainfrom
docs-numba-thread-tuning
Sep 22, 2026
Merged

TomaSusi merged 2 commits into
mainfrom
docs-numba-thread-tuning

Conversation

@pzeiger

@pzeiger pzeiger commented Sep 22, 2026 •

Copy link
Copy Markdown
Member

Summary

The "Set internal thread parallelization" section already covers oversubscription between Dask's chunk-level parallelism and FFTW/MKL's own per-transform threading. The same risk applies to Numba: the real-space multislice algorithm, and a few other CPU-only code paths, use Numba's own thread pool rather than the fftw.threads/mkl.threads knobs already documented, controlled instead by OMP_NUM_THREADS (or NUMBA_NUM_THREADS directly, which takes precedence) since abTEM/abTEM#441.

Adds one paragraph to that same section rather than a new one, since it's the same underlying principle stated for a different library.

What it says

  • Numba's thread pool follows OMP_NUM_THREADS if set, or every visible core otherwise; NUMBA_NUM_THREADS overrides independently.
  • These kernels are memory-bandwidth-bound and scale poorly past a handful of threads even for one isolated computation — a large value rarely helps, so benchmark rather than assuming the full core count is best.
  • Once Dask is already running many chunks concurrently, keeping the per-kernel thread count low costs nothing at worst and can measurably help — once the worker count exceeds the physical core count, extra internal threads per kernel add contention rather than finding anything left to parallelize into.

🤖 Written by Claude Code — Paul reviewed and posted it

pzeiger and others added 2 commits September 22, 2026 06:07
The real-space multislice algorithm and a few other CPU-only code
paths use Numba's own thread pool, controlled by OMP_NUM_THREADS (or
NUMBA_NUM_THREADS directly) rather than the fftw.threads/mkl.threads
knobs already documented here, but subject to the same oversubscription
risk against Dask's own chunk-level parallelism. Measured: these
kernels are memory-bandwidth-bound and scale poorly past a handful of
threads even for one isolated computation, and the per-kernel thread
count stops mattering altogether once Dask is running many chunks
concurrently.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Confirmed independently on a second machine (10-core/20-thread Xeon,
20 concurrent Dask-like callers): when worker count exceeds physical
core count, capping the per-kernel thread count measured 15-18%
faster, not merely indistinguishable, since a lower-core-count-per-
worker-count machine leaves no spare capacity for a kernel's own
internal threads to exploit even in principle.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
@TomaSusi
TomaSusi merged commit 0993949 into main Sep 22, 2026
1 check passed
@TomaSusi
TomaSusi deleted the docs-numba-thread-tuning branch September 22, 2026 06:39
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants