Conversation
The number of instances of each universe was found by searching the geometry from the root once per universe, so initialization took time proportional to the number of universes times the size of the geometry. Models with many universes, such as TRISO particles in a fine lattice, spent many minutes in this step. The universes are now ordered so that each comes after the universes containing it, and the counts are propagated from the root in that order, visiting each universe once. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014RN7JroEkrbkLKVHbMDap9
Every fill cell and lattice tile stored one offset per distribcell map, where a map is a universe containing distributed cells. By default every universe containing a material cell is a map, so the tables grew with the number of such universes times the size of the geometry. In models with many universes, such as TRISO particles in a fine lattice where each tile has its own universe, they could take tens of gigabytes. An offset is only read when the map's universe is in the cell's fill or the tile's universe. Each universe and lattice now keeps the sorted list of maps it contains, and each fill cell and tile stores offsets for those maps only. Maps are numbered in post-order, so the maps in a universe are mostly consecutive and a lookup is usually a single range check. The offsets are found in one visit of each universe and lattice from the root, instead of one pass over the geometry per map. Instance numbers are unchanged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014RN7JroEkrbkLKVHbMDap9
04dedba to
ac50777
Compare
|
@nuclearkevin this PR came out of trying the full 160 cm HTGR compact from your delta tracking work (the one with the TRISO search lattice) on With this PR, on my reconstruction of your compact (single thread, default settings):
Instance numbers are unchanged, and transport on the compact is the same or slightly faster. |
|
Thanks for looking into this @GuySten! I wonder if this will also decrease the time spent constructing the majorant in the delta tracking branch for TRISO problems, which has to do a forall cell instances lookup to find the maximum cell density per material. I'll try this out and see how it performs there. Edit: this also reminds me, I need to look at the 2-component combined estimator PR you put in. I'll get to that this weekend, sorry about the wait! |
Description
Two steps of geometry setup scale with the number of universes times the size of the geometry:
Models with many universes hit both. One example is a 160 cm TRISO fuel compact where each tile of a TRISO search lattice has its own universe: 20,485 maps and 239,490 fill cells. On
develop, the offset tables for this model need 21.3 GB (19.6 GB for cells, 1.7 GB for the lattice), while the rest of the simulation needs under 400 MB. Withmaterial_cell_offsets = False, the model fits in memory, but setup takes 19 minutes, almost all of it counting instances.This PR has one commit for each step:
DistribcellOffsetsclass. The offsets of a universe's fill cells, and of a lattice's tiles, are stored together in one array. Maps are numbered in post-order, so the maps in a universe are mostly consecutive and a lookup is usually one range check, with a binary search as fallback. The offsets are found in one visit of each universe and lattice from the root, instead of one pass over the geometry per map.The stored values are the same as before, so instance numbers don't change. Callers of
Cell::offset_andLattice::offset()are unchanged.distribcell_path_innernow skips cells and tiles that don't contain the map; they contain no instances of it, so the path it finds is the same.Results
160 cm TRISO compact, default settings, single thread:
developdevelop,material_cell_offsets = Falsek-effective is identical. On a 10 cm slice of the same compact, setup goes from 11–13.5 s to 6.8 s.
Transport cost
Share of profiled run time in the instance lookup (
cell_instance_at_leveland the offset accessors), single thread, two or three runs each:developOn PWR, wall-clock transport time over 6 alternating runs is within noise of
develop(medians differ by 0.9%, fastest runs by −2.4%).Testing
developon 17 geometries: rect and hex lattices, outer universes, nested lattices, a synthetic model with shared universes at several depths, TRISO, PWR and ATR.slice_datacell instances at every geometry level andDistribcellFilterbins, on slices along three axes: 265,791 distinct (cell, instance) pairs, all identical.tallies.outfrom short runs withDistribcellFiltertallies on 13 models: identical todevelop, including about 10,000 distributed cell labels.tests/unit_tests/test_universe_instances.pyfor the instance counts.develop.Checklist
I have made corresponding changes to the documentation (if applicable)