Repository navigation
Document Numba/OMP_NUM_THREADS oversubscription alongside FFT threading - #29
Merged
Merged
Conversation
The real-space multislice algorithm and a few other CPU-only code paths use Numba's own thread pool, controlled by OMP_NUM_THREADS (or NUMBA_NUM_THREADS directly) rather than the fftw.threads/mkl.threads knobs already documented here, but subject to the same oversubscription risk against Dask's own chunk-level parallelism. Measured: these kernels are memory-bandwidth-bound and scale poorly past a handful of threads even for one isolated computation, and the per-kernel thread count stops mattering altogether once Dask is running many chunks concurrently. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Confirmed independently on a second machine (10-core/20-thread Xeon, 20 concurrent Dask-like callers): when worker count exceeds physical core count, capping the per-kernel thread count measured 15-18% faster, not merely indistinguishable, since a lower-core-count-per- worker-count machine leaves no spare capacity for a kernel's own internal threads to exploit even in principle. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
The "Set internal thread parallelization" section already covers oversubscription between Dask's chunk-level parallelism and FFTW/MKL's own per-transform threading. The same risk applies to Numba: the real-space multislice algorithm, and a few other CPU-only code paths, use Numba's own thread pool rather than the
fftw.threads/mkl.threadsknobs already documented, controlled instead byOMP_NUM_THREADS(orNUMBA_NUM_THREADSdirectly, which takes precedence) since abTEM/abTEM#441.Adds one paragraph to that same section rather than a new one, since it's the same underlying principle stated for a different library.
What it says
OMP_NUM_THREADSif set, or every visible core otherwise;NUMBA_NUM_THREADSoverrides independently.🤖 Written by Claude Code — Paul reviewed and posted it