Problem
Croqtile's liveness-driven DMA resource allocator colors completion handles
(DMA futures and events) by live range into a minimal set of slots, so handles
with non-overlapping live ranges can share a slot.
On the CUDA backend, a scalar event that is "ready" can be lowered to a
hardware named barrier (sm_90+). Named barriers are a scarce, fixed pool; when
the pool is exhausted, the backend falls back to a scalar mbarrier.
Event arrays (event[N]) currently bypass this coloring. Liveness analysis
keys an ElemOf reference (ev[i]) by its base symbol, so the whole array is
treated as a single handle. Named-barrier lowering rejects multi-element
arrays and always lowers them to a single mbarrier array. As a result,
per-element slot reuse is lost even when the elements have pairwise disjoint
live ranges and could fit the named-barrier pool.
Proposal
Scalar-replace event arrays so each element becomes an independent completion
handle. Two routes are possible (both compiler-side):
Route A - virtual per-element naming in liveness
- Keep the array in the AST.
- Give each element its own virtual scoped name and live range
(ev[0]..ev[N-1]).
- Codegen lowers
ev[i] to barrier[base + i], where base is the array's
colored starting slot.
- Handles runtime-indexed
ev[i].
- Requires per-element lowering in codegen.
Route B - AST-level scalar replacement
- A scalar-replacement pass rewrites
event[N] into N real scalar symbols
(ev0..evN-1), rewriting the declaration and every constant-indexed
ElemOf reference into a distinct symbol.
- Downstream liveness and codegen see real scalar events and need no change.
- Only constant indices can be split; runtime
ev[i] needs guards or falls
back to the array.
Why not a size-N contiguous block
Modeling an array as a contiguous size-N resource would reserve N named-barrier
slots that no codegen consumes (arrays fall back to mbarrier). It would inflate
the color count and could starve scalar events out of the named-barrier pool.
Per-element unit resources are what both coloring and codegen can actually use.
Acceptance criteria
- An N-element event array whose elements have pairwise disjoint live ranges
can share slots with other events.
- Codegen emits N named barriers (or mbarriers) indexed by element slot.
- Existing end-to-end tests keep passing; add a test exercising array-element
slot reuse.
Problem
Croqtile's liveness-driven DMA resource allocator colors completion handles
(DMA futures and events) by live range into a minimal set of slots, so handles
with non-overlapping live ranges can share a slot.
On the CUDA backend, a scalar event that is "ready" can be lowered to a
hardware named barrier (sm_90+). Named barriers are a scarce, fixed pool; when
the pool is exhausted, the backend falls back to a scalar mbarrier.
Event arrays (
event[N]) currently bypass this coloring. Liveness analysiskeys an
ElemOfreference (ev[i]) by its base symbol, so the whole array istreated as a single handle. Named-barrier lowering rejects multi-element
arrays and always lowers them to a single mbarrier array. As a result,
per-element slot reuse is lost even when the elements have pairwise disjoint
live ranges and could fit the named-barrier pool.
Proposal
Scalar-replace event arrays so each element becomes an independent completion
handle. Two routes are possible (both compiler-side):
Route A - virtual per-element naming in liveness
(
ev[0]..ev[N-1]).ev[i]tobarrier[base + i], wherebaseis the array'scolored starting slot.
ev[i].Route B - AST-level scalar replacement
event[N]into N real scalar symbols(
ev0..evN-1), rewriting the declaration and every constant-indexedElemOfreference into a distinct symbol.ev[i]needs guards or fallsback to the array.
Why not a size-N contiguous block
Modeling an array as a contiguous size-N resource would reserve N named-barrier
slots that no codegen consumes (arrays fall back to mbarrier). It would inflate
the color count and could starve scalar events out of the named-barrier pool.
Per-element unit resources are what both coloring and codegen can actually use.
Acceptance criteria
can share slots with other events.
slot reuse.