Repository navigation
fix(analysis): count work and traffic per executing scope - #199
Merged
Merged
Conversation
zhen8838
commented
Oct 2, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
Fixes #194. Implicit loop captures made invariant values inherit their first use's scope, corrupting lifetime and repetition counts; whole-value movement also omitted executing mesh positions.
What
Contract
Numeric transitions are separated by capture correction, mesh positions, invariant loops and footprint pricing. Two transitions for a field use the measured M1 intermediate. Logical accounts omit invariant repetition; executed totals include every outer loop. Performance counts loop trips without serializing parallel mesh units. Footprint logical remains a single window; Call totals count waves and Function totals union within scope before multiplying waves and loop trips. Roofline consumes those totals once, falling back to traffic when footprint is unavailable or at another level. Programs, tutorial prose and notebook metadata are unchanged.
Capture scope correction
Mesh positions
PersistentGemmFlat memory.traffic.gmem.total.read 244.00MB -> 31.45GB (Mesh positions)
PersistentGemmFlat memory.traffic.gmem.total.write 280.00MB -> 36.09GB (Mesh positions)
PersistentGemmFlat memory.traffic.rmem.total.read 1.26GB -> 166.57GB (Mesh positions)
PersistentGemmFlat memory.traffic.rmem.total.write 1.26GB -> 104.67GB (Mesh positions)
PersistentGemmFlat memory.traffic.smem.total.read 1.17GB -> 46.41GB (Mesh positions)
PersistentGemmFlat memory.traffic.smem.total.write 1.17GB -> 61.88GB (Mesh positions)
GEMM_8192X17408X5120_OPTIMAL memory.traffic.gmem.total.read 1.01GB -> 133.68GB (Mesh positions)
GEMM_8192X17408X5120_OPTIMAL memory.traffic.gmem.total.write 306.00MB -> 39.45GB (Mesh positions)
GEMM_8192X17408X5120_OPTIMAL memory.traffic.rmem.total.read 2.71GB -> 357.29GB (Mesh positions)
GEMM_8192X17408X5120_OPTIMAL memory.traffic.rmem.total.write 2.71GB -> 357.20GB (Mesh positions)
GEMM_8192X17408X5120_OPTIMAL memory.traffic.smem.total.read 24.97MB -> 3.22GB (Mesh positions)
GEMM_8192X17408X5120_OPTIMAL memory.traffic.smem.total.write 24.97MB -> 3.22GB (Mesh positions)
Stage2[ctx=1816] memory.traffic.gmem.total.read 5.31MB -> 30.03MB (Mesh positions)
Stage2[ctx=1816] memory.traffic.gmem.total.write 3.55MB -> 28.39MB (Mesh positions)
Stage2[ctx=1820] memory.traffic.gmem.total.read 5.31MB -> 30.07MB (Mesh positions)
Stage2[ctx=1820] memory.traffic.gmem.total.write 3.56MB -> 28.45MB (Mesh positions)
Stage2[ctx=128] memory.traffic.gmem.total.read 1.60MB -> 11.90MB (Mesh positions)
Stage2[ctx=128] memory.traffic.gmem.total.write 258.62KB -> 2.02MB (Mesh positions)
Stage3[ctx=4096] memory.traffic.gmem.total.read 3.32MB -> 10.21MB (Mesh positions)
Stage3[ctx=4096] memory.traffic.gmem.total.write 4.00MB -> 4.02MB (Mesh positions)
Stage3[ctx=4096] memory.traffic.rmem.total.read 768B -> 24.00KB (Mesh positions)
Stage3[ctx=4096] memory.traffic.rmem.total.write 128B -> 4.00KB (Mesh positions)
Stage3[ctx=4096] memory.traffic.smem.total.read 5.57MB -> 20.83MB (Mesh positions)
Stage3[ctx=4096] memory.traffic.smem.total.write 5.40MB -> 20.53MB (Mesh positions)
Stage4[ctx=4096] memory.traffic.gmem.total.read 26.95MB -> 213.38MB (Mesh positions)
Stage4[ctx=4096] memory.traffic.gmem.total.write 24.38MB -> 195.05MB (Mesh positions)
Stage4[ctx=4096] memory.traffic.smem.total.read 323.25KB -> 337.25KB (Mesh positions)
Stage4[ctx=4096] memory.traffic.smem.total.write 322.25KB -> 329.25KB (Mesh positions)
Stage4[ctx=4096]/Call[v2] memory.traffic.smem.total.read 128.50KB -> 132.00KB (Mesh positions)
Stage4[ctx=4096]/Call[v43] memory.traffic.smem.total.read 128.50KB -> 132.00KB (Mesh positions)
Stage5[ctx=4096] memory.traffic.gmem.total.read 7.32MB -> 22.20MB (Mesh positions)
Stage5[ctx=4096] memory.traffic.gmem.total.write 4.00MB -> 32.02MB (Mesh positions)
Stage5[ctx=4096] memory.traffic.rmem.total.read 2.50KB -> 20.00KB (Mesh positions)
Stage5[ctx=4096]/Call[v21] memory.traffic.rmem.total.read 32B -> 256B (Mesh positions)
wgmma_cta_grid_4x17 memory.traffic.gmem.total.read 18.00KB -> 1.20MB (Mesh positions)
wgmma_cta_grid_4x17 memory.traffic.rmem.total.read 40.12KB -> 2.66MB (Mesh positions)
wgmma_cta_grid_4x17 memory.traffic.rmem.total.write 44.00KB -> 2.92MB (Mesh positions)
wgmma_cta_grid_4x17 memory.traffic.smem.total.read 1.12KB -> 76.50KB (Mesh positions)
wgmma_cta_grid_4x17 memory.traffic.smem.total.write 1.12KB -> 76.50KB (Mesh positions)
wgmma_cast_between_schedules memory.traffic.gmem.total.read 16.00KB -> 14.00KB (Mesh positions)
wgmma_cast_between_schedules memory.traffic.rmem.total.read 62.81KB -> 59.81KB (Mesh positions)
wgmma_cast_between_schedules memory.traffic.rmem.total.write 54.88KB -> 51.88KB (Mesh positions)
sm80_mma_ldmatrix memory.traffic.rmem.total.read 3.06KB -> 1.56KB (Mesh positions)
sm80_mma_ldmatrix memory.traffic.rmem.total.write 3.25KB -> 1.94KB (Mesh positions)
type_printer_sugar/Call[v6] memory.traffic.gmem.total.read 64B -> 256B (Mesh positions)
type_printer_sugar/Call[v6] memory.traffic.rmem.total.write 64B -> 256B (Mesh positions)
Qwen.layer_decode[128,128] memory.gmem.total 106862100 -> 13678348800 (Mesh positions)
Qwen.model[512,512] memory.gmem.total 37847437056 -> 4844471943168 (Mesh positions)
TiledQKVProjection memory.gmem.total.read 142606336 -> 18824036352 (Mesh positions)
TiledQKVProjection memory.gmem.total.write 16777216 -> 2214592512 (Mesh positions)
GemmWaveReuse memory.gmem.total.read 3244032 -> 428212224 (Mesh positions)
WithScope compute-cost.flops.f32.total 128 -> 256 (Mesh positions)
UnshardedInScope compute-cost.flops.f32.total 512 -> 1024 (Mesh positions)
InvariantReuse memory.gmem.total.read 192 -> 768 (Mesh positions)
MoEMegaKernel.experts memory.gmem.total r120.00KB/w90.00KB -> r7.79MB/w3.93MB (Mesh positions)
Invariant loops
Footprint pricing
Stage0[ctx=128] roofline.ideal_ns 632 -> 630 (Footprint pricing)
Stage2[ctx=1820] roofline.ideal_ns 1939 -> 1938 (Footprint pricing)
Stage2[ctx=128] roofline.ideal_ns 405 -> 404 (Footprint pricing)
Stage4[ctx=4096] roofline.ideal_ns 11213 -> 11159 (Footprint pricing)
Stage5[ctx=4096] roofline.ideal_ns 2474 -> 2473 (Footprint pricing)
Stage0[ctx=512] roofline.ideal_ns 1656 -> 1649 (Footprint pricing)
Stage0[ctx=1024] roofline.ideal_ns 3022 -> 3007 (Footprint pricing)
Stage0[ctx=2048] roofline.ideal_ns 5752 -> 5724 (Footprint pricing)
Stage0[ctx=4096] roofline.ideal_ns 11214 -> 11159 (Footprint pricing)
Stage0[ctx=8192] roofline.ideal_ns 22136 -> 22027 (Footprint pricing)
InvariantReuse memory.footprint.total 16 -> 192 (Footprint pricing)
WaveTruncation memory.footprint.total 1056 -> 2112 (Footprint pricing)
TruncatedWaveReuse memory.footprint.total 32 -> 64 (Footprint pricing)
WithScope memory.footprint.total 1024 -> 2048 (Footprint pricing)
UnshardedInScope memory.footprint.total 512 -> 1024 (Footprint pricing)
Qwen.layer_decode[128,128] memory.gmem.logical 102172180 -> 102172180 (unchanged; final total/logical=133.875472)
Qwen.model[512,512] memory.gmem.logical 13352177408 -> 13352177408 (unchanged; final total/logical=362.822617)
Risk
Roofline omits L2 until the target publishes calibrated L2 bandwidth.
compute_nsdoes not round for a tail wave. Function footprint groups are summed across scopes without cross-group deduplication, a pessimistic approximation; partial tail waves count as full waves. The existing first-solution allocator can exceed the aligned live-byte upper bound when corrected captures stay live across a loop (LEFTOVERS #14); allocation is unchanged. Same-wave reuse remains optimistic, and lower-level windows can count shared parent data again.