Repository navigation
fix(schedule): support broadcasts on symbolic meshes - #208
Merged
zhen8838 merged 19 commits intoOct 4, 2026
Merged
Conversation
Remove the unexecuted transfer fallback after schedule, TIR, analysis, passes, and HIR coverage all showed zero hits. Relation errors indicate invalid relations and must remain visible. Different-shape shard Copy verification is a separate contract outside this fix.
Unify scalar and tensor patterns while retaining ranked predicates on the bare Tensor singleton. Binary and Unary inputs accept rank-zero tensors, and placement-polymorphic UMAT values satisfy concrete selection storage requirements. Materialize scalar operands once per instruction through the shared explicit and automatic lowering entry. Paired HIR/TIR fixtures record AllocTensor and Fill, while the existing elementwise launch verifies scalar broadcasting on CUDA.
Build both the HIR site and candidate instruction relation from the projected operand types. Comparing global HIR coordinates with local instruction coordinates incorrectly filters singleton axes before selection can record candidates or refusals.
Keep the chunk-parallel RMS normalization fixture symbolic through parsing, then bind its closed envelope for analysis and candidate selection. Cover the legal upper bound, out-of-range and missing bindings, seven reshard sites, and five automatic binary candidates without changing its literal schedule operands. Factor dynamic multi-axis shards with successive exact structural quotients, sharing the local-shape helper, and recognize equal symbolic extents as one unsplit tile. Preserve integer validation and reject undecidable canonical divisions. Validate size-one lhs lowering in the existing scalar HIR/TIR golden pair; the plain fixture remains an analysis and selection program because its reshard instructions are intentionally unselected.
zhen8838
commented
Oct 3, 2026
Replace enumerated isl position sets with CuTe filter, directional coalescing, and continuity. Continuous selections compare their starts and normalized sizes so both column-major ldmatrix layouts and row-major symbolic meshes keep the same scope. Symbolic containment uses dim_range through dim_at_most; unproved bounds remain refused. Remove the three test cases for symbolic-position enumeration and affine-set errors because that representation and its 256-position limit no longer exist. Retain the four behavior cases and source-location wrappers. Profile filtering shares the existing coalesce traversal; staged type imports remain lazy to avoid the layout_algebra -> utils -> mesh initialization cycle.
Make is_contiguous consume the already filtered, coalesced input and state that precondition in its docstring and spec. Both scope callers now filter each arrangement only once and pass the same major parameter. Import size from its defining module and explain the staged core/types and utils-to-mesh import cycles in the owning docstrings.
Port the general composition, logical_divide and zipped_divide rules alongside the existing swizzle dispatch. Express col/row traversal through major parameters, retain per-mode tuple dispatch, and gather tile/remainder modes with the CuTe hierarchy rules. Symbolic tiles remain unsupported. The independent oracle records 96 comparisons against installed torch pycute: 16 input/tile pairs, three operations and both orders, including its documented composition examples and the three D17 forms. Do not route mesh _nested through divide: grouping axes by topology size is a different operation. Keep check_topology on layout.size; tile_view_layout is wired in the following commit.
Project the validated schedule repeat through each operand relation, and pass repeat from lowering into MmaAtom.operand_tiles. Tile views receive those counts instead of recomputing ratios from whole and tile shapes; preserve the grouped and flat prefix traversal, Split remapping, and diagnostics. Remove _tile_counts and its symbolic shape-equality exception. Do not replace prefix extraction with zipped_divide: its tile-first result does not describe the existing group-first flat fragments. The symbolically shaped mean-plus-EPS schedule has repeat (1,1,1), tensor counts (1,1,1), and scalar counts (); its layouts pass without extent division. Thirty-one captured real tile views stay identical, and ir/schedule/parser/ops-tir verification passes (602 passed, six existing diagnostic-fixture skips).
Expose program_dim_vars and update its CLI and corpus callers. Keep require_bound_dims and suggested_extents in CLI source: they assemble shared --dim guidance, not analysis knowledge. Move the unchanged structural quotient body to dim.exact_quotient. Shard layout callers defer importing it because core.op loads the staged types exports before core.expr defines Call; document that initialization cycle and the existing utils-to-shard_layout guard. Do not reorder bootstrap. Both CLI failure messages and no-report behavior remain byte-identical; cli/ir/schedule/analysis verification passes 629 tests.
Replace Scalar and Tensor pattern singletons with is_scalar_tensor and is_ranked_tensor in predicates.py. Move Ranked onto Predicate.holds and use the existing matcher dispatch; retain its declaration and refusal printer rules. Keep tensor_in unchanged. Migrate runtime declarations, the ParamDef example and existing test consumers without adding tests. Compare every one of the 2556 ranked-baseline and 180 scalar-baseline subject matches against the tile-ai#208 snapshots: all remain identical, including unrestricted HIR Binary and Unary slots.
Replace the four branch cases with a scope-comparison table and a gaps-refusal table, each with an unconditional assertion body. Retain both gap selections refusing their own scope and the full scope, and the make_mesh continuity diagnostic. The first table adds mirror assertions on the same input pairs: seven facts become ten, without changing input combinations. Before filling the table, probe within(full, flat), covered(full, half), and covered(full, THR); they evaluate to true, false, and false. Do not restore both_symbolic, enumeration_limit, or non_affine cases removed in M0: their ISL position-enumeration and affine-rendering constraints disappeared with level_positions. Scope continuity and containment remain public contracts.
Add offset * 0.5 after the automatic lhs-scalar subtraction and feed its result to the existing explicit size-one-lhs multiply. Regenerate the TIR golden: explicit 0.25 and automatic 1.0/0.5 operands each allocate and fill rank-zero register tensors. Explain why explicit scheduling with a literal as its first operand is a different path. The HIR register peak stays 132 bytes: offset dies when scaled replaces its tile, and the umat literal materializes only during lowering. Update only scalar_binary ledger derivation. Its rmem traffic changes from r396/w516 to r528/w644, which witnesses the new operation; do not alter memory or liveness semantics.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
What
4c38804a, source 1,373 passed / 10 existing skips (163s); isolated installed wheel: 162 passed / 0 skips (297s).Contract
DimVarbounds are inclusive[lo, hi]; migrated declarations retain their value sets.TensorPatternaccepts any rank;ScalarPatternand bareScalar/Tensoraliases are removed.is_scalar_tensor()constructs rank-zero patterns;is_ranked_tensor()uses publicRanked(Predicate)to preserve ranked matching. All 2,556 declaration matches and 180 scalar matches are unchanged by this naming refactor.umatfor concrete storage requirements; lowering materializes scalar operands withAllocTensor/Fillbefore TIR.T.Binaryverification and CUDA emission broadcast lhs/rhs symmetrically and preserve operand order.level_positionsis replaced byfilterand continuity checks; ISL position enumeration and its 256 limit disappear. Scope continuity remains mandatory, including symbolic containment checks.layout_algebraadds genericcomposition,filter,is_contiguous,logical_divide, andzipped_divide, all withmajor="col"|"row".is_contiguousconsumes a filtered layout; divide requires static integer tiles.tile_view_layout/tile_inner_typetake caller-suppliedcounts;operand_tilesprojects the supplied validatedrepeatinto operand counts._tile_countsis deleted. Shape-divisibility validation becomes layout-prefix validation; their coverage differs. The callers' repeat/inferred checks establish shape consistency.schedule candidatesaccepts repeatable--dim; missing/out-of-range bindings produce no partial report.canonical_shard_layoutsupports symbolic multi-axis structural division and rejects undecidable quotients.exact_quotientnow belongs todim.py;program_dim_varsis public. Quotient imports remain deferred because of staged package initialization.Risk
tests/analysis: 73.55s versus 70.77s before fix(schedule): support broadcasts on symbolic meshes #208 (+3.9%) and 74.27s in fix(schedule): support broadcasts on symbolic meshes #208 (-1.0%). Repeated measurements vary by approximately 0.35s; these comparisons include that noise. Single-scope measurements changed from 213/431µs to 47.6/92.1µs (one/two levels).