Skip to content
Merged
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
12 changes: 12 additions & 0 deletions docs/user_guide/appendix/performance_tips.ipynb
Original file line number Diff line number Diff line change
Expand Up @@ -278,6 +278,18 @@
"number of workers times the FFT threads at or below your physical core count. FFTW also exposes its planner through\n",
"`fftw.planning_effort`; more patient planning can pay off for long-running simulations that reuse one grid size.\n",
"\n",
"The same principle applies to Numba, which a few CPU-only code paths use directly — the\n",
"real-space multislice algorithm among them. Numba's thread pool follows `OMP_NUM_THREADS` if it\n",
"is set, or every visible core otherwise, and `NUMBA_NUM_THREADS` overrides that independently of\n",
"the other libraries above. These kernels are memory-bandwidth-bound and scale poorly past a\n",
"handful of threads even in isolation, so a large value rarely helps a single computation —\n",
"benchmark before assuming your full core count is best. Once Dask is already running many chunks\n",
"concurrently, keeping this low costs nothing at worst, and can measurably help: when the number\n",
"of concurrent workers already exceeds your physical core count — common once a worker pool is\n",
"sized to logical rather than physical cores — each kernel call claiming its own extra threads adds\n",
"contention with nothing left to actually parallelize into, and can cost real throughput rather\n",
"than merely fail to gain any.\n",
"\n",
"### Use \"good\" numbers of `gpts`\n",
"\n",
"FFT implementations are most efficient for array sizes that factor into small primes (2, 3, 5 and 7). Since the exact\n",
Expand Down
Loading