Conversation
prn() converted the 64-bit output of the generator to a double in [0, 1) with ldexp(result, -64), which compiles to an out-of-line library call on every random number. Multiplying by the exact power of two 0x1p-64 gives bit-for-bit identical results (the scaling is exact and cannot underflow) and compiles to a single instruction. Random numbers and simulation results are unchanged. Transport time is reduced by about 7% for a neutron source in a heavy water sphere with thermal scattering and by about 4% for a gamma shielding problem. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014RN7JroEkrbkLKVHbMDap9
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
prn()converts the 64-bit output of the PCG generator to a double in [0, 1) withldexp(result, -64). This compiles to an out-of-line library call on every random number, and profiling shows it is a noticeable fraction of run time (scalbn/__scalbntogether were ~7% of a heavy water problem with thermal scattering and ~5% of a gamma shielding problem).This PR replaces it with a multiplication by the exact power of two
0x1p-64. Because the conversion fromuint64_ttodoubleis unchanged and scaling by a power of two is exact (the result is either zero or at least 2⁻⁶⁴, so it cannot underflow), the result is bit-for-bit identical toldexp. I checked this on 2×10⁸ inputs, including edge cases such as 0, 2⁵³±1 and 2⁶⁴−1, with no mismatches. The now-unused<cmath>include is removed.Random number sequences and simulation results are unchanged; the full unit and regression test suite gives the same results as
develop, and tallies were bit-for-bit identical in the timing runs below.Transport time, single thread, average of three runs:
Checklist
I have followed the style guidelines for Python source files (if applicable)I have made corresponding changes to the documentation (if applicable)I have added tests that prove my fix is effective or that my feature works (if applicable)