Skip to content

fix(schedule): support broadcasts on symbolic meshes - #208

Merged
zhen8838 merged 19 commits into
tile-ai:mainfrom
zhen8838:fix/schedule-broadcast-symbolic-mesh
Oct 4, 2026
Merged

zhen8838 merged 19 commits into
tile-ai:mainfrom
zhen8838:fix/schedule-broadcast-symbolic-mesh

Conversation

@zhen8838

@zhen8838 zhen8838 commented Oct 3, 2026 •

Copy link
Copy Markdown
Collaborator

Why

What

  • Share broadcast-aware relations, compare candidate relations on local tiles, and materialize scalar operands through the common explicit/automatic lowering entry.
  • Replace custom scope machinery with CuTe layout algebra, pass validated repeat counts to tile views, and reorganize pattern and dimension helpers.
  • Cover symbolic RMS normalization, scalar operands on both sides, lhs broadcasting, and lowering goldens through existing workflows.
  • Frozen verification: 4c38804a, source 1,373 passed / 10 existing skips (163s); isolated installed wheel: 162 passed / 0 skips (297s).

Contract

  • DimVar bounds are inclusive [lo, hi]; migrated declarations retain their value sets.
  • TensorPattern accepts any rank; ScalarPattern and bare Scalar/Tensor aliases are removed. is_scalar_tensor() constructs rank-zero patterns; is_ranked_tensor() uses public Ranked(Predicate) to preserve ranked matching. All 2,556 declaration matches and 180 scalar matches are unchanged by this naming refactor.
  • Selection accepts umat for concrete storage requirements; lowering materializes scalar operands with AllocTensor/Fill before TIR.
  • T.Binary verification and CUDA emission broadcast lhs/rhs symmetrically and preserve operand order.
  • level_positions is replaced by filter and continuity checks; ISL position enumeration and its 256 limit disappear. Scope continuity remains mandatory, including symbolic containment checks.
  • layout_algebra adds generic composition, filter, is_contiguous, logical_divide, and zipped_divide, all with major="col"|"row". is_contiguous consumes a filtered layout; divide requires static integer tiles.
  • tile_view_layout/tile_inner_type take caller-supplied counts; operand_tiles projects the supplied validated repeat into operand counts. _tile_counts is deleted. Shape-divisibility validation becomes layout-prefix validation; their coverage differs. The callers' repeat/inferred checks establish shape consistency.
  • schedule candidates accepts repeatable --dim; missing/out-of-range bindings produce no partial report.
  • canonical_shard_layout supports symbolic multi-axis structural division and rejects undecidable quotients. exact_quotient now belongs to dim.py; program_dim_vars is public. Quotient imports remain deferred because of staged package initialization.

Risk

Remove the unexecuted transfer fallback after schedule, TIR, analysis, passes, and HIR coverage all showed zero hits. Relation errors indicate invalid relations and must remain visible. Different-shape shard Copy verification is a separate contract outside this fix.
Unify scalar and tensor patterns while retaining ranked predicates on the bare Tensor singleton. Binary and Unary inputs accept rank-zero tensors, and placement-polymorphic UMAT values satisfy concrete selection storage requirements.

Materialize scalar operands once per instruction through the shared explicit and automatic lowering entry. Paired HIR/TIR fixtures record AllocTensor and Fill, while the existing elementwise launch verifies scalar broadcasting on CUDA.
Build both the HIR site and candidate instruction relation from the projected operand types. Comparing global HIR coordinates with local instruction coordinates incorrectly filters singleton axes before selection can record candidates or refusals.
Keep the chunk-parallel RMS normalization fixture symbolic through parsing, then bind its closed envelope for analysis and candidate selection. Cover the legal upper bound, out-of-range and missing bindings, seven reshard sites, and five automatic binary candidates without changing its literal schedule operands.

Factor dynamic multi-axis shards with successive exact structural quotients, sharing the local-shape helper, and recognize equal symbolic extents as one unsplit tile. Preserve integer validation and reject undecidable canonical divisions.

Validate size-one lhs lowering in the existing scalar HIR/TIR golden pair; the plain fixture remains an analysis and selection program because its reshard instructions are intentionally unselected.
Comment thread src/tilefoundry/cli/source.py
Comment thread src/tilefoundry/ir/pattern/pattern.py Outdated
Comment thread src/tilefoundry/ir/pattern/pattern.py Outdated
Comment thread src/tilefoundry/ir/types/mesh.py Outdated
Comment thread src/tilefoundry/ir/types/shard_layout.py Outdated
Comment thread tests/fixtures/schedule/hir/scalar_binary.py
Comment thread tests/ir/types/test_mesh.py Outdated
Replace enumerated isl position sets with CuTe filter, directional coalescing, and continuity. Continuous selections compare their starts and normalized sizes so both column-major ldmatrix layouts and row-major symbolic meshes keep the same scope. Symbolic containment uses dim_range through dim_at_most; unproved bounds remain refused.

Remove the three test cases for symbolic-position enumeration and affine-set errors because that representation and its 256-position limit no longer exist. Retain the four behavior cases and source-location wrappers. Profile filtering shares the existing coalesce traversal; staged type imports remain lazy to avoid the layout_algebra -> utils -> mesh initialization cycle.
Make is_contiguous consume the already filtered, coalesced input and state that precondition in its docstring and spec. Both scope callers now filter each arrangement only once and pass the same major parameter. Import size from its defining module and explain the staged core/types and utils-to-mesh import cycles in the owning docstrings.
Port the general composition, logical_divide and zipped_divide rules alongside the existing swizzle dispatch. Express col/row traversal through major parameters, retain per-mode tuple dispatch, and gather tile/remainder modes with the CuTe hierarchy rules. Symbolic tiles remain unsupported.

The independent oracle records 96 comparisons against installed torch pycute: 16 input/tile pairs, three operations and both orders, including its documented composition examples and the three D17 forms. Do not route mesh _nested through divide: grouping axes by topology size is a different operation. Keep check_topology on layout.size; tile_view_layout is wired in the following commit.
Project the validated schedule repeat through each operand relation, and pass repeat from lowering into MmaAtom.operand_tiles. Tile views receive those counts instead of recomputing ratios from whole and tile shapes; preserve the grouped and flat prefix traversal, Split remapping, and diagnostics. Remove _tile_counts and its symbolic shape-equality exception.

Do not replace prefix extraction with zipped_divide: its tile-first result does not describe the existing group-first flat fragments. The symbolically shaped mean-plus-EPS schedule has repeat (1,1,1), tensor counts (1,1,1), and scalar counts (); its layouts pass without extent division. Thirty-one captured real tile views stay identical, and ir/schedule/parser/ops-tir verification passes (602 passed, six existing diagnostic-fixture skips).
Expose program_dim_vars and update its CLI and corpus callers. Keep require_bound_dims and suggested_extents in CLI source: they assemble shared --dim guidance, not analysis knowledge.

Move the unchanged structural quotient body to dim.exact_quotient. Shard layout callers defer importing it because core.op loads the staged types exports before core.expr defines Call; document that initialization cycle and the existing utils-to-shard_layout guard. Do not reorder bootstrap. Both CLI failure messages and no-report behavior remain byte-identical; cli/ir/schedule/analysis verification passes 629 tests.
Replace Scalar and Tensor pattern singletons with is_scalar_tensor and is_ranked_tensor in predicates.py. Move Ranked onto Predicate.holds and use the existing matcher dispatch; retain its declaration and refusal printer rules. Keep tensor_in unchanged.

Migrate runtime declarations, the ParamDef example and existing test consumers without adding tests. Compare every one of the 2556 ranked-baseline and 180 scalar-baseline subject matches against the tile-ai#208 snapshots: all remain identical, including unrestricted HIR Binary and Unary slots.
Replace the four branch cases with a scope-comparison table and a gaps-refusal table, each with an unconditional assertion body. Retain both gap selections refusing their own scope and the full scope, and the make_mesh continuity diagnostic.

The first table adds mirror assertions on the same input pairs: seven facts become ten, without changing input combinations. Before filling the table, probe within(full, flat), covered(full, half), and covered(full, THR); they evaluate to true, false, and false.

Do not restore both_symbolic, enumeration_limit, or non_affine cases removed in M0: their ISL position-enumeration and affine-rendering constraints disappeared with level_positions. Scope continuity and containment remain public contracts.
Add offset * 0.5 after the automatic lhs-scalar subtraction and feed its result to the existing explicit size-one-lhs multiply. Regenerate the TIR golden: explicit 0.25 and automatic 1.0/0.5 operands each allocate and fill rank-zero register tensors. Explain why explicit scheduling with a literal as its first operand is a different path.

The HIR register peak stays 132 bytes: offset dies when scaled replaces its tile, and the umat literal materializes only during lowering. Update only scalar_binary ledger derivation. Its rmem traffic changes from r396/w516 to r528/w644, which witnesses the new operation; do not alter memory or liveness semantics.
@zhen8838
zhen8838 merged commit 21981d9 into tile-ai:main Oct 4, 2026
2 checks passed
@zhen8838
zhen8838 deleted the fix/schedule-broadcast-symbolic-mesh branch October 4, 2026 05:41
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fix(ir): handle a symbolic extent in selected_run fix(schedule): give an elementwise read the result's rank

1 participant