Skip to content

fix(analysis): count work and traffic per executing scope - #199

Merged
zhen8838 merged 10 commits into
tile-ai:mainfrom
zhen8838:fix/scope-multiplicity
Oct 2, 2026
Merged

zhen8838 merged 10 commits into
tile-ai:mainfrom
zhen8838:fix/scope-multiplicity

Conversation

@zhen8838

@zhen8838 zhen8838 commented Oct 1, 2026

Copy link
Copy Markdown
Collaborator

Why

Fixes #194. Implicit loop captures made invariant values inherit their first use's scope, corrupting lifetime and repetition counts; whole-value movement also omitted executing mesh positions.

What

  • Give loops explicit invariant captures and unify region bindings, isolation and variance with mesh regions; preserve canonical window arithmetic.
  • Share IterationScope repetition accounting across compute, memory and performance, and price roofline from executed footprint windows.
  • Preserve authored capture labels, regenerate analysis-only goldens/tutorial outputs, and verify all four tutorial pages from an installation.

Contract

Numeric transitions are separated by capture correction, mesh positions, invariant loops and footprint pricing. Two transitions for a field use the measured M1 intermediate. Logical accounts omit invariant repetition; executed totals include every outer loop. Performance counts loop trips without serializing parallel mesh units. Footprint logical remains a single window; Call totals count waves and Function totals union within scope before multiplying waves and loop trips. Roofline consumes those totals once, falling back to traffic when footprint is unavailable or at another level. Programs, tutorial prose and notebook metadata are unchanged.

Capture scope correction

  • Stage3[ctx=4096] memory.traffic.smem.total.read 5.60MB -> 5.57MB (Capture scope correction)
  • Stage3[ctx=4096] memory.traffic.smem.total.write 5.43MB -> 5.40MB (Capture scope correction)
  • Stage3[ctx=4096] memory.traffic.smem.cta.read 670.41KB -> 666.47KB (Capture scope correction)
  • Stage3[ctx=4096] memory.traffic.smem.cta.write 660.85KB -> 656.91KB (Capture scope correction)
  • Stage3[ctx=4096] roofline.ideal_ns 1600 -> 1599 (Capture scope correction)
  • Stage5[ctx=4096] memory.traffic.smem.total.read 20.87MB -> 20.78MB (Capture scope correction)
  • Stage5[ctx=4096] memory.traffic.smem.total.write 20.55MB -> 20.49MB (Capture scope correction)
  • Stage5[ctx=4096] memory.traffic.smem.cta.read 2.61MB -> 2.60MB (Capture scope correction)
  • Stage5[ctx=4096] memory.traffic.smem.cta.write 2.57MB -> 2.56MB (Capture scope correction)
  • Stage5[ctx=4096] memory.peak.smem 40.94KB -> 40.81KB (Capture scope correction)
  • wgmma_one_tile_of_larger_output memory.peak.gmem 38.00KB -> 22.00KB (Capture scope correction)
  • gqa_decode.GqaOnline.gqa_online_attend[ctx_len=128] memory.peak.gmem 284672 -> 282624 (Capture scope correction)
  • prefill_decode_attention.PrefillDecodeAttention.attend[ctx=128,seq=128] memory.peak.smem 278528 -> 245760 (Capture scope correction)
  • qwen3_1_7b_pd.PrefillLayer.layer_prefill[ctx_len=128,seq=128] memory.peak.gmem 163193872 -> 171582480 (Capture scope correction)
  • qwen3_1_7b_pd.PrefillLayer.layer_prefill[ctx_len=128,seq=128] memory.peak.rmem 198144 -> 132608 (Capture scope correction)
  • qwen3_1_7b_pd.PrefillLayer.model[ctx_len=0,seq=512] memory.peak.gmem 5304603152 -> 5269475856 (Capture scope correction)
  • qwen3_1_7b_pd.PrefillLayer.model[ctx_len=0,seq=512] memory.peak.rmem 263680 -> 132608 (Capture scope correction)
  • qwen3_1_7b_pd.PrefillLayer.model[ctx_len=512,seq=512] memory.peak.gmem 5304603152 -> 5269475856 (Capture scope correction)
  • qwen3_1_7b_pd.PrefillLayer.model[ctx_len=512,seq=512] memory.peak.rmem 263680 -> 132608 (Capture scope correction)
  • qwen3_1_7b_pd.PrefillLayer.model[ctx_len=512,seq=1] memory.peak.gmem 4602883616 -> 4611717152 (Capture scope correction)
  • qwen3_1_7b_pd.PrefillLayer.model[ctx_len=512,seq=1] memory.peak.rmem 1644 -> 1036 (Capture scope correction)
  • qwen3_1_7b_pd.PrefillLayer.model[ctx_len=4608,seq=1] memory.peak.gmem 4602883616 -> 4611717152 (Capture scope correction)
  • qwen3_1_7b_pd.PrefillLayer.model[ctx_len=4608,seq=1] memory.peak.rmem 1644 -> 1036 (Capture scope correction)
  • Qwen.layer_decode[128,128] memory.peak.gmem 145374224 -> 154289168 (Capture scope correction)
  • Qwen.layer_decode[128,128] memory.peak.rmem 1132 -> 1036 (Capture scope correction)
  • Qwen.layer_decode[128,128] memory.gmem.total 166901662134272 -> 106862100 (Capture scope correction)
  • Qwen.model[512,512] memory.gmem.total 39540225229035090605506560 -> 37847437056 (Capture scope correction)
  • Stage3[ctx=4096] memory.footprint.repeat_interleave_k.logical 2.00MB -> 2.06MB (Capture scope correction)
  • Stage3[ctx=4096] memory.footprint.repeat_interleave_v.logical 2.00MB -> 2.06MB (Capture scope correction)
  • Stage5[ctx=4096] memory.footprint.cache_update_k.logical 128B -> 512.12KB (Capture scope correction)
  • Stage5[ctx=4096] memory.footprint.cache_update_v.logical 128B -> 512.12KB (Capture scope correction)
  • Stage5[ctx=4096] memory.footprint.cur_pos.combined_logical 8B -> 4B (Capture scope correction)
  • Stage3[ctx=4096] memory.reuse.q_rope.holds_bytes 128.50KB -> absent (Capture scope correction)
  • Stage3[ctx=4096] memory.reuse.q_rope.reuse_bytes 3.50KB -> absent (Capture scope correction)
  • Stage3[ctx=4096] memory.footprint.alias.k_heads.logical 64.00KB -> merged (Capture scope correction)
  • Stage3[ctx=4096] memory.footprint.alias.v_heads.logical 64.00KB -> merged (Capture scope correction)
  • Stage5[ctx=4096] memory.footprint.alias.k_all.logical 512.00KB -> merged (Capture scope correction)
  • Stage5[ctx=4096] memory.footprint.alias.v_all.logical 512.00KB -> merged (Capture scope correction)
  • Stage5[ctx=4096] memory.footprint.alias.cur_pos_2.logical 4B -> merged (Capture scope correction)

Mesh positions

  • PersistentGemmFlat memory.traffic.gmem.total.read 244.00MB -> 31.45GB (Mesh positions)

  • PersistentGemmFlat memory.traffic.gmem.total.write 280.00MB -> 36.09GB (Mesh positions)

  • PersistentGemmFlat memory.traffic.rmem.total.read 1.26GB -> 166.57GB (Mesh positions)

  • PersistentGemmFlat memory.traffic.rmem.total.write 1.26GB -> 104.67GB (Mesh positions)

  • PersistentGemmFlat memory.traffic.smem.total.read 1.17GB -> 46.41GB (Mesh positions)

  • PersistentGemmFlat memory.traffic.smem.total.write 1.17GB -> 61.88GB (Mesh positions)

  • GEMM_8192X17408X5120_OPTIMAL memory.traffic.gmem.total.read 1.01GB -> 133.68GB (Mesh positions)

  • GEMM_8192X17408X5120_OPTIMAL memory.traffic.gmem.total.write 306.00MB -> 39.45GB (Mesh positions)

  • GEMM_8192X17408X5120_OPTIMAL memory.traffic.rmem.total.read 2.71GB -> 357.29GB (Mesh positions)

  • GEMM_8192X17408X5120_OPTIMAL memory.traffic.rmem.total.write 2.71GB -> 357.20GB (Mesh positions)

  • GEMM_8192X17408X5120_OPTIMAL memory.traffic.smem.total.read 24.97MB -> 3.22GB (Mesh positions)

  • GEMM_8192X17408X5120_OPTIMAL memory.traffic.smem.total.write 24.97MB -> 3.22GB (Mesh positions)

  • Stage2[ctx=1816] memory.traffic.gmem.total.read 5.31MB -> 30.03MB (Mesh positions)

  • Stage2[ctx=1816] memory.traffic.gmem.total.write 3.55MB -> 28.39MB (Mesh positions)

  • Stage2[ctx=1820] memory.traffic.gmem.total.read 5.31MB -> 30.07MB (Mesh positions)

  • Stage2[ctx=1820] memory.traffic.gmem.total.write 3.56MB -> 28.45MB (Mesh positions)

  • Stage2[ctx=128] memory.traffic.gmem.total.read 1.60MB -> 11.90MB (Mesh positions)

  • Stage2[ctx=128] memory.traffic.gmem.total.write 258.62KB -> 2.02MB (Mesh positions)

  • Stage3[ctx=4096] memory.traffic.gmem.total.read 3.32MB -> 10.21MB (Mesh positions)

  • Stage3[ctx=4096] memory.traffic.gmem.total.write 4.00MB -> 4.02MB (Mesh positions)

  • Stage3[ctx=4096] memory.traffic.rmem.total.read 768B -> 24.00KB (Mesh positions)

  • Stage3[ctx=4096] memory.traffic.rmem.total.write 128B -> 4.00KB (Mesh positions)

  • Stage3[ctx=4096] memory.traffic.smem.total.read 5.57MB -> 20.83MB (Mesh positions)

  • Stage3[ctx=4096] memory.traffic.smem.total.write 5.40MB -> 20.53MB (Mesh positions)

  • Stage4[ctx=4096] memory.traffic.gmem.total.read 26.95MB -> 213.38MB (Mesh positions)

  • Stage4[ctx=4096] memory.traffic.gmem.total.write 24.38MB -> 195.05MB (Mesh positions)

  • Stage4[ctx=4096] memory.traffic.smem.total.read 323.25KB -> 337.25KB (Mesh positions)

  • Stage4[ctx=4096] memory.traffic.smem.total.write 322.25KB -> 329.25KB (Mesh positions)

  • Stage4[ctx=4096]/Call[v2] memory.traffic.smem.total.read 128.50KB -> 132.00KB (Mesh positions)

  • Stage4[ctx=4096]/Call[v43] memory.traffic.smem.total.read 128.50KB -> 132.00KB (Mesh positions)

  • Stage5[ctx=4096] memory.traffic.gmem.total.read 7.32MB -> 22.20MB (Mesh positions)

  • Stage5[ctx=4096] memory.traffic.gmem.total.write 4.00MB -> 32.02MB (Mesh positions)

  • Stage5[ctx=4096] memory.traffic.rmem.total.read 2.50KB -> 20.00KB (Mesh positions)

  • Stage5[ctx=4096]/Call[v21] memory.traffic.rmem.total.read 32B -> 256B (Mesh positions)

  • wgmma_cta_grid_4x17 memory.traffic.gmem.total.read 18.00KB -> 1.20MB (Mesh positions)

  • wgmma_cta_grid_4x17 memory.traffic.rmem.total.read 40.12KB -> 2.66MB (Mesh positions)

  • wgmma_cta_grid_4x17 memory.traffic.rmem.total.write 44.00KB -> 2.92MB (Mesh positions)

  • wgmma_cta_grid_4x17 memory.traffic.smem.total.read 1.12KB -> 76.50KB (Mesh positions)

  • wgmma_cta_grid_4x17 memory.traffic.smem.total.write 1.12KB -> 76.50KB (Mesh positions)

  • wgmma_cast_between_schedules memory.traffic.gmem.total.read 16.00KB -> 14.00KB (Mesh positions)

  • wgmma_cast_between_schedules memory.traffic.rmem.total.read 62.81KB -> 59.81KB (Mesh positions)

  • wgmma_cast_between_schedules memory.traffic.rmem.total.write 54.88KB -> 51.88KB (Mesh positions)

  • sm80_mma_ldmatrix memory.traffic.rmem.total.read 3.06KB -> 1.56KB (Mesh positions)

  • sm80_mma_ldmatrix memory.traffic.rmem.total.write 3.25KB -> 1.94KB (Mesh positions)

  • type_printer_sugar/Call[v6] memory.traffic.gmem.total.read 64B -> 256B (Mesh positions)

  • type_printer_sugar/Call[v6] memory.traffic.rmem.total.write 64B -> 256B (Mesh positions)

  • Qwen.layer_decode[128,128] memory.gmem.total 106862100 -> 13678348800 (Mesh positions)

  • Qwen.model[512,512] memory.gmem.total 37847437056 -> 4844471943168 (Mesh positions)

  • TiledQKVProjection memory.gmem.total.read 142606336 -> 18824036352 (Mesh positions)

  • TiledQKVProjection memory.gmem.total.write 16777216 -> 2214592512 (Mesh positions)

  • GemmWaveReuse memory.gmem.total.read 3244032 -> 428212224 (Mesh positions)

  • WithScope compute-cost.flops.f32.total 128 -> 256 (Mesh positions)

  • UnshardedInScope compute-cost.flops.f32.total 512 -> 1024 (Mesh positions)

  • InvariantReuse memory.gmem.total.read 192 -> 768 (Mesh positions)

  • MoEMegaKernel.experts memory.gmem.total r120.00KB/w90.00KB -> r7.79MB/w3.93MB (Mesh positions)

Invariant loops

  • PersistentGemmFlat compute-cost.other_ops.integer.total 8580 -> 16896 (Invariant loops)
  • PersistentGemmFlat compute-cost.other_ops.integer.cta 65 -> 128 (Invariant loops)
  • PersistentGemmFlat compute-cost.other_ops.integer.thread 65 -> 128 (Invariant loops)
  • PersistentGemmFlat performance.predicted_ns 13392246 -> 17938389 (Invariant loops)
  • Stage3[ctx=4096] compute-cost.other_ops.integer.total 288 -> 512 (Invariant loops)
  • Stage3[ctx=4096] compute-cost.other_ops.integer.cta 9 -> 16 (Invariant loops)
  • SymbolicStoreOffset compute-cost.integer.per_cta.delta 6 -> 192 (Invariant loops)
  • SymbolicStoreOffset compute-cost.integer.total.delta 768 -> 24576 (Invariant loops)
  • SymbolicStoreOffset performance.predicted_ns.delta 6 -> 192 (Invariant loops)
  • Literal/SymbolicStoreOffset roofline.ideal_ns 398459 -> 418219 (Invariant loops)

Footprint pricing

  • Stage0[ctx=128] roofline.ideal_ns 632 -> 630 (Footprint pricing)

  • Stage2[ctx=1820] roofline.ideal_ns 1939 -> 1938 (Footprint pricing)

  • Stage2[ctx=128] roofline.ideal_ns 405 -> 404 (Footprint pricing)

  • Stage4[ctx=4096] roofline.ideal_ns 11213 -> 11159 (Footprint pricing)

  • Stage5[ctx=4096] roofline.ideal_ns 2474 -> 2473 (Footprint pricing)

  • Stage0[ctx=512] roofline.ideal_ns 1656 -> 1649 (Footprint pricing)

  • Stage0[ctx=1024] roofline.ideal_ns 3022 -> 3007 (Footprint pricing)

  • Stage0[ctx=2048] roofline.ideal_ns 5752 -> 5724 (Footprint pricing)

  • Stage0[ctx=4096] roofline.ideal_ns 11214 -> 11159 (Footprint pricing)

  • Stage0[ctx=8192] roofline.ideal_ns 22136 -> 22027 (Footprint pricing)

  • InvariantReuse memory.footprint.total 16 -> 192 (Footprint pricing)

  • WaveTruncation memory.footprint.total 1056 -> 2112 (Footprint pricing)

  • TruncatedWaveReuse memory.footprint.total 32 -> 64 (Footprint pricing)

  • WithScope memory.footprint.total 1024 -> 2048 (Footprint pricing)

  • UnshardedInScope memory.footprint.total 512 -> 1024 (Footprint pricing)

  • Qwen.layer_decode[128,128] memory.gmem.logical 102172180 -> 102172180 (unchanged; final total/logical=133.875472)

  • Qwen.model[512,512] memory.gmem.logical 13352177408 -> 13352177408 (unchanged; final total/logical=362.822617)

Risk

Roofline omits L2 until the target publishes calibrated L2 bandwidth. compute_ns does not round for a tail wave. Function footprint groups are summed across scopes without cross-group deduplication, a pessimistic approximation; partial tail waves count as full waves. The existing first-solution allocator can exceed the aligned live-byte upper bound when corrected captures stay live across a loop (LEFTOVERS #14); allocation is unchanged. Same-wave reuse remains optimistic, and lower-level windows can count shared parent data again.

Comment thread src/tilefoundry/analysis/loop_domain.py Outdated
Comment thread src/tilefoundry/analysis/roofline.py
Comment thread tests/inspection/test_roundtrip.py Outdated
Comment thread tests/parser/test_slices.py Outdated
@zhen8838
zhen8838 merged commit 0c82146 into tile-ai:main Oct 2, 2026
1 check passed
@zhen8838
zhen8838 deleted the fix/scope-multiplicity branch October 2, 2026 05:47
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fix(analysis): charge a replicated tile once per mesh position

1 participant