Skip to content

Guided cleanup: naming, one plan-time barrier, restructured tests and benchmarks - #22

Merged
lkdvos merged 60 commits into
mainfrom
cleanup
Oct 7, 2026
Merged

lkdvos merged 60 commits into
mainfrom
cleanup

Conversation

@lkdvos

@lkdvos lkdvos commented Oct 7, 2026 •

Copy link
Copy Markdown
Member

A chunk-by-chunk review of the whole package, aimed at a codebase that is easier to read and maintain. Each chunk was explained, discussed and then changed in its own commit(s), verified with the full suite, bitwise old-vs-new equivalence scripts where behaviour had to stay the same, per-call floor and allocation checks, TTFX and specialisation counts. Performance claims come from same-node Slurm A/B runs.

Breaking changes (0.1, unregistered)

  • Descriptive M/N/K names throughout, including the public API: plan_contract's mc/kc/nc keywords and the Blocking fields are now m_block/k_block/n_block; mr(k)/nr(k) are tile_size(k) / tile_size(k, i) (now public, with sliver_width).
  • plan_contract and contract! now also reject an output that might alias an input, as the TensorOperations backend already did (coarse for now: any shared parent; see Precise aliasing check between output and inputs #18).

What changed, by area

  • Naming and layout. No leading underscores on internal names (constants included); files organised by concern; Holy traits with the docstring on the abstract type.
  • Hardware (target.jl). Own detection kept; TargetProfile derives vector width, register count, per-core cache shares and the cache-line size at construction.
  • Axis groups and tiles. One AxisGroup constructor from labels; one-based coordinates inside tiles, zero-based addresses; one Tile type with getindex/setindex!; ScatterAxis is pointer-backed so the axis union stays unboxed.
  • Packing. One pack!(panel, tile, sliver_spec(kernel, i), transform) with A and B on equal footing (B packed as transpose(tile)); line packing in packing/line_packing.jl; UnpackedBView is a B source next to PackedPanel.
  • Microkernels. The internal ComplexMethod traits are gone: the unparameterised kernel type (e.g. PlanarKernel) names a scheme. A kernel is a K step plus an accumulator layout (real, split, lane pairs); one generic add_tile/store_tile!; β cases dispatch on VectorInterface's Zero/One at the existing branch points, with the β hoist kept. Unpacked-B codegen fix: dense B loads through per-column pointers and no unrolling of that K loop (ComplexF64 small M 2× faster on Icelake, Float64 small M/64³ 16–27% faster on Genoa, Rome neutral).
  • Planning. One select_shape pipeline; analytical blocking with the hybrid complex rule (k_block from the real default kernel's B sliver, m_block/n_block from the kernel's own bytes); no workspace pool and no tile-wise oracle (reuse a plan to reuse buffers).
  • One dispatch barrier, at plan time. Blocking, line packing, the C panel and the execution path are decided before the kernel barrier; the path is part of the plan's type, so execute! on a reused plan has no dynamic dispatch. The hint machinery, storage stripping and the workspace's slot cache are gone. The static ramp flags stay: a runtime-ramp variant was measured (same runtime, but TTFX +40–90% and the suite 22 → 32 min) and reverted.
  • Execution. execute.jl split into execute.jl, nest.jl and c_panel.jl; trivial EmptyPath/ScalePath for empty extents and K = 0; one tile loop for packed and unpacked B; the dot path dispatches on its orientation and the workspace is sized per path (dot plans allocate 8.5 KB instead of 150 KB).
  • TensorOperations backend. No continuation struct: α/β travel in the plan request; contract! goes through one dynamic crossing instead of two (2×2: 530 → 455 ns).
  • Tests. ParallelTestRunner (each file in its own process and module), self-contained files mirroring src/, per-folder helpers, Aqua ambiguity checks for the package itself. A coverage audit drove trims that lose no src line and keep every eltype × kernel family and path flag reaching the nest. Pkg.test 22 min → under 3 min (--check-bounds=yes kept; julia_args = ["--check-bounds=auto"] is the fast local run).
  • Benchmarks. Answered experiments removed; probes for TTFX, call floor and specialisation counts; same-node A/B script (submit_ab.sh); --smoke everywhere; benchmark/README.md; a new suite overview plot (one row per eltype: QuasiStrided GF/s, StridedBLAS GF/s, time ratio, against arithmetic intensity, coloured by total work, with tag filters).

Follow-ups

🤖 Generated with Claude Code

lkdvos and others added 30 commits September 30, 2026 15:18
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Decision D1: names follow `<m|n|k>_<level>_<what>` codebase-wide.

Public renames:
- `Blocking` fields `mc, kc, nc` -> `m_block, k_block, n_block`
- `plan_contract` keywords `mc/kc/nc` -> `m_block/k_block/n_block`
- `mr(kernel)`, `nr(kernel)` -> `tile_size(kernel) -> (m_tile, n_tile)`
- `packed_a_per_k`, `packed_b_per_k` -> `sliver_widths(kernel)`
  (`tile_size` and `sliver_widths` are now in the `public` list)

`using LinearAlgebra` is dropped from src: it used no LinearAlgebra names
(`adjoint`/`transpose` are Base's).

Full test suite: 34542/34542 pass. TTFX unchanged (Float64 first call
3.3-3.5 s before, 3.2-3.7 s after).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
`sliver_widths` becomes `sliver_width`, matching `size`; both accessors gain
an index method, replacing the `tile_size(k)[i]` spelling.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- The ISA is the CPUID probe's; the LLVM CPU-name table is gone (it could
  only narrow the probe, e.g. capping AVX-512 Zen 4/5 at AVX2 on Julia 1.10).
- `CacheLevel` drops the unused `ways`; `TargetProfile` drops `arch` and
  derives `vector_bytes`/`nregisters` from `isa`, rejecting unknown ISAs.
- `core_bytes(profile, level)` replaces three copies of the per-core cache
  share in blocking.jl, pack_split.jl and labels.jl.
- No leading underscores in target.jl; `_TARGET` is `TARGET`.
- Tests: synthetic profiles take only the ISA, so the sweeps over impossible
  (isa, vector_bytes, nregisters) combinations go (34542 -> 33444 tests);
  the forced-ISA runner takes QS_FAKE_ISA without QS_FAKE_VB/QS_FAKE_NREG.

Full suite passes (15 min).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- `pair_group.jl` becomes the `AxisGroup(labels, (ind1, v1), (ind2, v2))`
  constructor.
- `normalize_group` is gone and `offsets` moves to the test helpers (both
  were test-only).
- `map_ramp_step` moves from dot.jl; `affine_ramp` is built on it.
- `Base.setindex`, `prod` and `Base.Checked` replace hand-written versions.
- `BlockDescriptor`/`describe_block` move to tiles.jl; `describe_block` checks
  `count <= length(buffer) - first`, so `first + count` cannot overflow past
  the check.
- `fill_offsets!` drops its alias check.
- No leading underscores in axis_group.jl.

Full suite passes (30989 tests, 15 min; the drop is the removed
normalize_group/offsets/alias tests).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- AffineAxis <: AbstractVector{Int} (one-based getindex, closed-form
  checked extrema); the scatter axis is a plain view of the offset buffer.
  ScatterAxis, PtrScatterAxis, axis_from_descriptor and the Axis union go.
- Tile{S, R, C} replaces QSTile/SourceTile/DestinationTile; size and
  @propagate_inbounds getindex/setindex! replace nrows/ncols/tile_offset/
  tile_load/tile_store!, with @inbounds at the former call sites.
- extrema replaces axis_offset_range; is_unit_stride (with an
  AbstractVector fallback) replaces _unit_stride_rows.
- Tile coordinates and axis_offset stay zero-based for now (temporary shims).
- The nest passes axes to their consumers through _with_axis, one branch
  per axis type: a Union holding a view is heap-boxed (432 B to 1.6 KB per
  call on a contraction with scattered M, N and K). A zero-allocation test
  for that contraction is added.

Tests: 30930 pass (30989 before; removed ScatterAxis/PtrScatterAxis/
axis_from_descriptor tests).
TTFX probe, Float64 / ComplexF64 first call, interleaved runs:
  before 3.1-3.3 s / 2.1-2.3 s, after 3.4-3.7 s / 2.3-2.4 s.
Reused-plan execute! medians (us, 3 interleaved runs, before -> after):
  matmul 256^3 Float64 404 -> 395, ComplexF64 1537 -> 1442;
  scattered-M 12x20x256x256 Float64 432 -> 431, ComplexF64 1460 -> 1456;
  scattered-M/N/K small Float64 17.2 -> 17.6, ComplexF64 40.4 -> 39.1.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Tile indexing, AffineAxis/view axes, packer lane and K-step loops,
  microkernel K steps, columns and row lanes (generated code included) and
  UnpackedBView columns/K steps are one-based; the two zero-based shims go.
- packed_{a,b}_offset and packed_{a,b}_plane_offset take one-based
  coordinates and still return zero-based offsets: the panel layout is
  unchanged. Addresses, block starts and AxisGroup coordinates stay
  zero-based, as does the macro-block coordinate of the line-by-line packer.

Tests: 30930 pass.
TTFX probe, Float64 / ComplexF64 first call, interleaved with the chunk-2
commit: before 3.1-3.2 s / 2.1-2.2 s, after 3.3-3.4 s / 2.2-2.4 s.
Reused-plan execute! medians (us, 3 interleaved runs, before -> after):
  matmul 256^3 Float64 411 -> 400, ComplexF64 1564 -> 1532;
  scattered-M 12x20x256x256 Float64 448 -> 431, ComplexF64 1503 -> 1499;
  scattered-M/N/K small Float64 17.9 -> 17.6, ComplexF64 42.0 -> 41.4.
bench_kernels.jl kernel arm, Float64, median of 5 interleaved runs (GF/s):
  16x6/W8 85.8 -> 84.7, 32x6/W8 107.0 -> 106.9.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A `view` scatter axis is not `isbits`, so the nest's
`Union{AffineAxis, view}` was heap-boxed, which `_with_axis` worked around
with about ten nested closures. `ScatterAxis` is again a borrowed pointer,
now an `AbstractVector{Int}` indexed one-based, so `_axis_of` returns an
unboxed union and the nest reads flat. The oracle and `_scale_all_of_C!`
now `GC.@preserve` the workspace they borrow from.

Full suite passes (30930 tests). Reused-plan execute! medians, closures ->
pointer, us: matmul 256^3 F64 410 -> 410, C64 1578 -> 1559; scattered M F64
432 -> 441, C64 1545 -> 1531; scattered M/N/K F64 18.4 -> 18.1, C64 41.9 ->
42.0. Zero allocations per call in all six, under --check-bounds=yes. First
call F64/C64: 3.5-3.8 s / 2.4-2.8 s -> 3.2 s / 2.1-2.2 s (back to the
pre-chunk-3 3.1-3.3 s / 2.1-2.3 s).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Panels are only `PackedPanel`: the `AbstractVector` methods of
  `panel_vload`/`panel_load`/`panel_store!` are gone. The checked entry
  points `pack_a!`, `pack_b!` and `accumulate` take a `DenseVector` and
  borrow it as a `PackedPanel` under `GC.@preserve`; `execute_tile!` and
  the oracle reach vectors through `accumulate`. The kernels' `accumulate`
  methods constrain the A panel to `PackedPanel` (B may be an
  `UnpackedBView`), which keeps the vector method unambiguous.
- One `Descriptor(Val(MR), Val(NR), T, a_format = RealFormat(),
  b_format = RealFormat())` constructor and a `RealDescriptor` alias replace
  `KernelDescriptor`/`ComplexKernelDescriptor`; `check_descriptor` names
  `Descriptor` in its messages.
- `packed_a_offset(d, i, p, plane = 0)` / `packed_b_offset(d, j, p,
  plane = 0)` replace the plane and real-only offset functions; the layout
  is unchanged. The packers' `(plane, i, p)` closures adapt the order.
- transposed.jl is merged into the end of panel.jl.
- Tests: view destinations dropped (no longer supported), the unsafe
  packer test borrows its buffers, the "complex kernel has no single-plane
  offset" MethodError test is replaced by a plane-offset check.

Measured against 81cfaa3 on a workstation:
- Tests: 30931 pass (baseline 30930).
- TTFX (load; first Float64, ComplexF64, K=1 outer call), 2 runs each,
  interleaved: before 0.5/2.9/2.0/0.6 s and 0.5/3.0/2.0/0.6 s, after
  0.5/3.0/2.1/0.6 s and 0.5/2.9/2.0/0.6 s.
- Reused-plan execute! medians over 5 interleaved processes (before ->
  after, us): matmul256 Float64 393.5 -> 391.8, ComplexF64 1472.7 ->
  1487.2; scatterM Float64 414.9 -> 414.5, ComplexF64 1445.8 -> 1460.7;
  scatterMN Float64 17.5 -> 16.7, ComplexF64 39.7 -> 39.2 (minima equal).
- 0 bytes allocated per call in every case under --check-bounds=yes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A and B are packed by one code path; the kernel's descriptor decides each
operand's packing through a `SliverSpec{I, L, F, T}` (side, lanes, format,
scalar type), from `sliver_spec(kernel, i)`.

- `pack!(panel, tile, spec, transform)` / `unsafe_pack!` replace
  pack_a!/pack_b!/unsafe_pack_a!/unsafe_pack_b! and the DescriptorKernel
  forwarding methods. Lanes run along `tile.rows`, K steps along `tile.cols`;
  B is packed as `transpose` of its tile (the nest and oracle pass B's axes
  swapped). `Base.transpose(::Tile)` added.
- `panel_offset(spec, t, p, plane)` is the one offset formula;
  `packed_a_offset`/`packed_b_offset`/`packed_*_length` are defined on it
  (layout unchanged).
- Fast paths are symmetric: real B with unit-stride columns now takes the
  `Vec{L}` load/convert/store path.
- `emit!` takes `(re, im, neg_im)`; the value and padding callers supply them
  (padding stays literal +0.0). The block packer takes the spec.
- One complex contiguous K loop with a per-format `store_step!` (LLVM IR of
  the shuffles/loads/stores identical to before for all formats x transforms).
- Dead `PackedPanel` eligibility conditions dropped; eltype error built in an
  `@noinline` helper; `is_unit_stride`, `dense_lanes`, `lane_convertible`,
  `complex_fastpath_isa_eligible` collected in pack_contiguous.jl; no leading
  underscores in pack.jl / pack_contiguous.jl.

Tests: 30931 -> 31537. Real A/B halves merged into sweeps over specs, plus
"B packs as A packs its transposed tile" and a real-B fast-path sweep
(L = 6); removed the duplicate ISA-gate testset (planar store), the
PackedPanel-gate eligibility checks, the fmaddsub Vector-vs-panel duplicate
and the unsafe_pack_b! duplicates.

TTFX (base / new, 2 interleaved runs): first call Float64 4.5, 3.4 / 4.2,
4.3 s; ComplexF64 3.1, 3.0 / 2.7, 2.7 s; load 0.5-0.6 s both.

A/B, reused-plan execute! medians (5 rounds, us, base -> new): matmul256
F64 513 -> 521, scatterM F64 600 -> 568, scatterMN F64 31.6 -> 32.0,
matmul256 C64 2005 -> 1981, scatterM C64 2011 -> 1963, scatterMN C64
62.1 -> 61.7. Row-major B (7 rounds): F64 571 -> 482, F32 282 -> 265
(5-round run: 605 -> 566, 232 -> 269; workstation noise is +-30%).
Scattered K (7 rounds): F64 582 -> 590, C64 2069 -> 1717. Isolated real B
sliver pack, fast vs scalar, kc = 256: 2-3x faster for every L in
4, 6, 8, 12, 16 (Float32 and Float64), so no power-of-two restriction.
Per-call allocations 0 for all cases under --check-bounds=yes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- `Base.accumulate` -> `add_tile`, a function of our own, in every kernel,
  the `DenseVector` method, callers and tests.
- `KernelMethod` (was `ComplexMethod`) singletons declare their packing
  formats once with `pack_formats`; `a_reals`/`b_reals` go, blocking uses
  `reals_per_element` of the formats. A test checks every kernel's descriptor
  agrees with `pack_formats(complex_method(k))`.
- `DescriptorKernel` -> `Microkernel`; leading underscores dropped from the
  names in interface.jl and scalar.jl.
- One `@inline execute_tile!` and one `@inline pack!`, whose storage check is
  `@boundscheck`; every other validation stays unconditional. The nest calls
  `@inbounds _pack_sliver!`/`@inbounds _execute_micro_tile!` (both
  `@propagate_inbounds`, replacing `unsafe_execute_micro_tile!` and the
  function-valued `pack!` argument); the oracle calls them checked. LLVM of
  `_execute_nest!`, `_micro_tiles_packed_b!` and `_micro_tiles_unpacked_b!`
  (Float64/ComplexF64, affine and scattered) has the same line count as before
  and no per-tile storage check; under `--check-bounds=yes` the per-tile
  checks now run, so the hoisting test counts them.
- `ScalarKernel` accumulates into an `NTuple{MR*NR,T}` (column-major) and its
  store branches on `beta` once per tile; it now allocates nothing.
- Contract moved from the header comment into docstrings (`Microkernel`,
  `KernelMethod`, `zero_accumulator`, `add_tile`, `store_tile!`,
  `execute_tile!`, `pack!`).

Tests: 31548 pass (was 31537).
TTFX (2 runs each, interleaved), before and after alike: load 0.5 s; first
call Float64 3.2-3.3 s, ComplexF64 2.3 s; K=1 outer 0.6 s.
A/B medians of 5 interleaved rounds, after/before: matmul256 F64 0.997,
C64 0.991; scatterM 1.017/0.994; scatterMN 1.003/1.001; rowmajB256 F64
1.009, F32 0.996; scatterK F64 0.996, C64 1.009. Allocations 0 per call
under --check-bounds=yes.
ScalarKernel 4x4 nest, 20x30x25 output, K = 16: Float64 155 -> 50 us,
Float32 159 -> 46-55 us; allocations per call 218400 / 151200 bytes -> 0.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
All vector kernel structs move to src/microkernels/types.jl under an
abstract VectorKernel{MR, NR, T, W}, with three accumulator layouts
(RealLayout, SplitLayout, LanePairLayout) and the generator-time pieces
of the stores (index, lane value, full-block store, store prelude), so
they dispatch instead of `<:` tests. One generic add_tile loops
accumulate_step over k_steps (1m: 2k); 1m and the mixed-domain kernels
forward their step to a computed inner(kernel) (no inner field, no KI
parameter; the call folds to a singleton). One generic store_tile!,
vector store and scalar store parameterised by the layout; @inline per
layout as before (real/split yes, lane pairs no). One generic
constructor with the complex-T check and vector-shape check, generic
lanewidth and zero_accumulator. 1m gets the lane-pair vector store.
Renames: b_scalar/b_complex (UnpackedBView hooks), accumulate_step,
vector_store!/scalar_store!/vector_store_eligible, split_index,
split_store_block!/lanepair_store_block!, fmaddsub/swap_pairs/addsub.

Codegen (AVX-512, default shapes): K-loop instruction counts identical
for all six kernels; store_tile!/vector/scalar store instruction counts
identical (SIMD 2015/1953/1317, Planar 1511/1494/907, FMAddSub
101/3619/2423, ComplexReal 101/3612/2417, RealComplex 1976/1900/1202);
1m now has the 3612-instruction lane-pair vector store.

add_tile is now @inline for every kernel (before: only SIMD and mixed).
A/B vs 95728d0, medians of 5 interleaved runs: small-K (M = N = 200,
K = 1..8) ComplexF64 -23..-35%, Float64 -3..+2%; 256^3 named kernels:
SIMD -1%, Planar -2.5%, 1m -11%, FMAddSub -7%, ComplexReal +3%,
RealComplex 0%; earlier matmul/scattered/row-major/scattered-K cases
-4..+2%. Allocations 0 under --check-bounds=yes.
TTFX: Float64 first call 3.1 s (before 3.1/3.4 s), ComplexF64 2.1 s
(before 2.1/2.2 s). Tests: 31967 pass (+419 for the 1m vector store),
suite ~18 min.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
vector_store! generates its block loop once per beta case (iszero,
isone, general) and branches once per tile; the full-block expressions
are the specialised ones (real: alpha*vec, muladd(alpha, vec, old),
muladd(alpha, vec, beta*old); split_store_block!/lanepair_store_block!
take the case as a Val, same arithmetic). Tails keep axpby_at!: inside
each arm LLVM folds its repeated beta test.

Codegen (AVX-512, default shapes): K loops unchanged. vector_store!
instructions before -> after: SIMD 1953 -> 2726, Planar 1494 -> 2004,
FMAddSub 3619 -> 4757, ComplexReal/1m 3612 -> 4485, RealComplex
1900 -> 2628. Beta compares in the vector store: SIMD 96 -> 2, Planar
62 -> 13, FMAddSub/ComplexReal 167 -> 28, RealComplex 83 -> 16. Code per
executed variant (beta 0/1/general): SIMD 1040/1112/1023, Planar
492/635/707, lane pairs ~1003/1291/1723, RealComplex 637/828/924.

A/B vs the previous commit, medians of 5 interleaved runs: small-K
(M = N = 200) Float64 K = 1: -12%, K = 2..8: -19..-27%; ComplexF64
-3..-10%; 256^3 matmul/scattered/named-kernel cases -4..+1%.
Allocations 0 under --check-bounds=yes.

Cost: TTFX first call Float64 3.1 -> 3.7 s, ComplexF64 2.0 -> 2.4 s
(2 runs each, consistent). Full suite 31967 pass, 28m02 (previous
commit 18m07; single runs on a shared workstation, but in line with the
TTFX increase).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Five files instead of eight, included in this order:

- interface.jl: the contract (Microkernel, KernelMethod and pack_formats,
  the add_tile/store_tile!/zero_accumulator docstrings, execute_tile!,
  scale_tile!, store_prologue!, axpby helpers, shape checks).
- kernels.jl: the seven kernel types and their type-level facts (method
  and layout traits, rows_per_vector, acc_index, inner, k_steps,
  lanewidth, the generic constructor). The 1m and mixed-domain header
  comments are folded into the kernels' docstrings.
- vecops.jl: fmaddsub, swap_pairs, addsub, deinterleave_re/_im,
  interleave_planes.
- steps.jl: the K steps (scalar, real, planar, fmaddsub), the B-element
  hooks, the generic add_tile and zero_accumulator.
- stores.jl: the layouts' generator-time pieces, the full-block stores and
  the generic store_tile!/vector_store!/scalar_store!, plus ScalarKernel's.

Removed: types.jl, scalar.jl, simd.jl, planar.jl, onem.jl, fmaddsub.jl,
mixed.jl.

Holy traits: accumulator_layout(kernel) becomes AccumulatorLayout(kernel)
and complex_method(kernel) becomes KernelMethod(kernel), each defined on
the kernel type, with the docstring on the abstract type. The
complex_method(::Any) fallback is replaced by KernelMethod for
ScalarKernel (its only user).

Every vector kernel's inner constructor now runs the vector-shape check
(check_vector_kernel), so K{MR, NR, T, W}(descriptor) is validated and the
generic constructor only builds the descriptor. A real T for a complex
kernel is rejected by the descriptor (its message instead of the
kernel's).

Behaviour-neutral: tests 31967/31967 pass (27m59s wall); K-loop and
store instruction counts identical to cc78fd1 for all six vector
kernels; TTFX unchanged (2 interleaved runs each, before/after: load
0.5/0.5-0.6 s, first Float64 call 3.7-3.8/3.7 s, ComplexF64 2.4/2.4 s,
K=1 outer 0.6/0.6 s).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- `_sort_perm` (hand-written insertion sort) replaced by
  `TupleTools.sortperm`/`getindices`. TupleTools sorts by a recursive merge
  that takes from the right half only on a strict `lt`, so ties keep their
  original order, as before; allocation-free and inferred for NTuples.
  `import TupleTools` moves to the module file.
- `_order_contract_labels` and `_choose_k_order` merged into
  `order_contract_labels` with an optional trailing `l2bytes` (default
  `nothing`: the core's L2 share, looked up only when a cost is needed).
- `_l2_core_bytes` (both methods) moved to defaults.jl as `l2_core_bytes`.
- conjugation.jl deleted: `_op_conjugates`/`_qs_isconj` now live in plan.jl
  as `op_conjugates`/`isconj(view, flag)`, both GUARDRAIL comments kept.
- Names defined in labels.jl lose their leading underscore (`classify_labels`,
  `order_free_labels`, `leading_unit_run`, `prefer_swap`, `KOperand`,
  `k_operand`, `k_order_cost`, ...). The `_K_*` constants keep theirs, matching
  the package's other internal constants. A test helper
  `_macro_op_conjugates` became `_macro_conj_op`.

Behaviour-neutral: K-order model, free-label order and swap rule unchanged.

Per-call floor (probe_call_floor.jl, 4 interleaved runs each, medians,
base 3be4a87 -> this commit, ns): F64 8x8x8 360 -> 369, 16^3 470 -> 472,
d4 outer 352 -> 365; C64 8x8x8 417 -> 417, 16^3 933 -> 931, d4 outer
576 -> 577. Sample ranges overlap (F64 d4 outer 347-389 vs 347-373);
0 B/call throughout. Planning micro-benchmark with 3 K labels:
K order 39 -> 18 ns, 0 B.

Tests: 31967/31967 pass, 28m52s.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
`select_shape(T, method, m_length, k_length, run)` in kernel_selection.jl
runs every automatic selection step in order (host shape, extent step-down,
small M, `fit_to_run`), previously split between `_default_shape` in
defaults.jl and `_plan_with_kernel` in plan.jl; one docstring lists the
steps. defaults.jl keeps only the per-eltype cache. The mixed-method mapping
now wraps the whole pipeline once. The swap decision calls it with an
unbounded `k_length`, so, as before, it judges orientations without the
K-dependent divisor step.

`fit_to_run` merges `_store_shape` and `_demote_shape_for_run`: the
MV = 4 -> MV = 2 step at any K, then the divisor step at small K, with their
conditions unchanged. It now runs after the small-M step; the store step
only fires on the tall shape, which the small-M step never leaves, so the
order is equivalent.

Also: the NEON planar override rows go (the register fit selects the same
shapes; a test pins them), the dead `vb % sizeof(real(T)) == 0` conditions
go, `_MixedMethod` -> `MixedMethod`, and names defined in kernel_selection.jl
lose their leading underscore. The uncached `_kernel_for` and the unused
`n_length` of the old pipeline are dropped.

Behaviour-neutral: the automatic kernel type, shape, blocking and
orientation are identical to 7032aed for Float32/64, ComplexF32/64 and both
mixed pairs, m, n, k in {1,2,3,5,8,13,17,31,48,64,100,257}, four C layouts,
under :avx512 (Intel and znver4), :avx2, :neon and :unknown profiles
(138270 plans), and on a dense (m <= 300, every run dividing m, 7 K values)
grid of the shape pipeline itself (579040 cases).

Per-call floor (medians of 3 interleaved runs, ns, before -> after):
Float64 8^3 372 -> 368, 16^3 484 -> 479, d4 outer 344 -> 340;
ComplexF64 8^3 424 -> 415, 16^3 966 -> 963, d4 outer 580 -> 589; 0 B.
TTFX: first call Float64 3.8 -> 3.8-3.9 s, ComplexF64 2.2-2.5 -> 2.4-2.5 s.

Tests: 32198 pass (was 31967; the mixed-selection grid gained a K
dimension), 48 min on a loaded workstation.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The unparameterised kernel type (`SIMDKernel`, `PlanarKernel`, ...) is now
the shape-free name of a complex-arithmetic scheme. `KernelMethod`, its six
singletons, `KernelMethod(kernel)` and `_kernel_type` are gone;
`default_kernel_type(T[, TA, TB])` replaces `default_method`. Menus, pack
formats, accumulator planes, blocking scaling, the shape rules, the mixed
mapping (`MixedKernel`), unpacked-B eligibility and the menu ladder dispatch
on `::Type{<:PlanarKernel}` etc.; a pair without a menu has the empty menu.
The per-scheme summary moves to the `pack_formats` docstring in kernels.jl.
ScalarKernel shares SIMDKernel's traits, so `default_blocking` works on it.
With eligibility predicted from the kernel type itself, `_hint_holds` loses
its eligibility clause.

Two places keep a type out of a value position, since a plain `Type`
regressed the per-call floor:
- the plan barrier receives `Val{<concrete kernel type>}()` (from
  `menu_val`) instead of `Val(shape)` plus the method: a `Type` argument
  misses the dynamic dispatch's fast cache (+350-430 ns per call measured
  with the bare type);
- `select_shape` returns the kernel type as a `Val`: inference widens a type
  in a returned tuple to `UnionAll`, which made `menu_val` a dynamic call
  that boxed the shape (32 B per complex plan); recovering the Union by an
  `isa` assertion or `===` on types cost a subtype check or structural
  egal (+30-100 ns).

Behaviour-neutral: the automatic kernel type, shape, blocking and
orientation are identical to 7032aed on the same 138270-plan grid as the
previous commit, and the shape pipeline on the dense 579040-case grid.

Per-call floor (medians of 3 interleaved runs vs 7032aed, ns): Float64 8^3
372 -> 369, 16^3 484 -> 474, d4 outer 352 -> 341; ComplexF64 8^3 426 -> 422,
16^3 957 -> 954, d4 outer 586 -> 575; 0 B. TTFX: first call Float64
3.7-3.8 -> 3.7 s, ComplexF64 2.4 -> 2.3 s.

Tests: 32190 pass (redundant `KernelMethod(k) === ...` checks removed, one
ScalarKernel blocking check added), 34 min.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Blocking: `default_blocking(kernel, profile = target_profile())` evaluates
the analytical model with the kernel's own tile and packed bytes per
element (B sliver = NR * packed reals * sizeof(real(T))), so the "same
packed byte budget" falls out of the model; undetected caches take the
fallback rows scaled by packed reals as before. `_scale_blocking`,
`_real_blocking_row`, the default-shape `_modelled_blocking` method and
`default_blocking(::Type)` are gone. Names in blocking.jl lose their
underscores (`modelled_blocking`, `fallback_blocking`, `kernel_blocking`,
`roundup`).

No host-defaults cache: `ResolvedDefaults`, the per-eltype slots and
defaults.jl are removed. The `Val(profile.isa)` tables become functions
branching on the `isa` Symbol; the small-M FMAddSub candidates are a filter
over the static menu. select_shape, blocking, the K-order model, the split
decision and the TensorOperations pool read `target_profile()` directly.
The direct version first cost +25-57 ns per Float64 call: the `cpu_name`
string compare in `profile_mv` (7 ns, three times per plan) and the
divisions of `core_bytes` in the blocking model. Passing the profile down
cannot help (`target_profile()` is a plain load), so instead of an
eltype-keyed cache the profile itself carries the three derived values,
computed by its constructor: `l2_share`, `l3_share` (the core's cache
shares) and `double_pumped` (znver4/znver5). The path hint computes a
kernel's default `k_block` only under a C narrower than T.

Line size: `line_bytes(profile)` (detected L1d line, else 64) replaces
`_K_LINE_BYTES` in the K-order model and the split decision;
`_K_WALK_FAR_BYTES`/`_K_WALK_FAR_PENALTY` move to target.jl.

Line packing: src/packing/line_packing.jl holds `PackSplit`, the split
decision, the split enumeration and the block packer, now
`pack_block_by_lines!(..., split::PackSplit)`.
`pack_split(g, kgroup, spec, S, k_block, eff, rounded, k_block_requested;
l2bytes = nothing)` reads the K map, tile and complexity from the
operand's `SliverSpec`. src/planning/pack_split.jl is deleted.

Test-only API removed: `_default_kernel(T, m, n)` and `_default_kernel(T)`;
tests use `select_shape` + `kernel_from_shape` (helpers `auto_kernel`,
`host_kernel`), the pool builds the host-shape kernel.

Behaviour changes (equivalence grid over 7 profiles, eltypes, extents,
C layouts, every menu kernel's blocking, line splits and K orders; the
dense select_shape grid is identical):
- Real kernels: no blocking changed on any profile. Mixed kernels: none.
- Complex planar/1m/FMAddSub kernels on detected caches block by their own
  tile. On this host (cascadelake): planar 24x3 unchanged (96,341,786);
  planar 16x6 192,170,1578; 1m 12x8 120,128,2096 (was 48,341,786);
  FMAddSub 12x8 252,128,2096 (was 96,341,786); the AVX2 planar 4x5
  160,204,1315 (was 96,341,786). Undetected caches: unchanged.
- 128-byte lines (an M1-like :neon profile, and :avx512 with line 128):
  K orders flip on 44 grid cases, and line splits change on 22 (new
  splits where a line is 8 instead of 4 elements; larger L chunks).
  64-byte profiles: no split or K-order decision changed.

Per-call floor (tensorcontract!, 6 interleaved rounds, medians, ns,
base -> new, 0 B both): Float64 8^3 362 -> 363, 16^3 473 -> 459,
d4 outer 362 -> 334; ComplexF64 8^3 412 -> 405, 16^3 967 -> 948,
d4 outer 570 -> 551. TTFX (3 rounds): Float64 3.8 -> 3.7 s,
ComplexF64 2.4 -> 2.3 s, K=1 outer 0.6 -> 0.6 s.

Tests: 32300 pass (was 32190; the per-kernel blocking test adds the rest),
54 min wall on the loaded workstation.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Same-node, interleaved per case, over every planar/1m/fmaddsub menu shape
that fits the host's vectors. Validates chunk 11's change of complex
blocking.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Workspace pooling is gone: the TensorOperations task-local pool, reserve!,
the reuse helpers and plan_contract's `workspace =` keyword. A plan always
builds its own workspace through the given allocator, and reusing the plan
is how buffers get reused. The two TO.tensorcontract! methods are now one:
checkpoint, then `_planned` with an `_Execute` continuation that executes
and releases the plan inside the kernel barrier (so nothing is boxed), then
reset. There is no try/finally, same as TO's own blas_contract!.

The oracle is gone too: execute_tilewise!, its tw_* buffers and the `oracle`
keyword. The tests that compared against it already compared execute!
against a dense matmul or `_lo_reference`, so those comparisons stay. The
ramp-vs-buffer test now checks against `_lo_reference`.

Workspace: the ten offset/descriptor fields become `m`/`n::GroupBuffers`
(offsets and descriptors, each a pair) and `k::NTuple{2, Vector{Int}}`. The
tile buffers are dropped; the beta-only pass uses the M/N offsets. Offsets
are plain Vectors whatever the allocator. There is one ContractWorkspace
constructor (tensoralloc for every allocator) plus `workspace_sizes`, and a
`release!(plan, allocator)` method. SlotCache/barrier_slot! move to
barrier.jl, which is now included before workspace.jl. The planning barrier
passes a fresh `Ref` because no workspace exists yet. ContractPlan documents
that it is used by one task at a time.

Tests: plan reuse with poisoned buffers is still exact, and so is the
beta-only pass. Removed the pool, reuse and reserve! tests. The TO
zero-byte-steady-state test is replaced by "the TO call allocates no more
than plan_contract" (the boxing guard). The plan_contract allocation ceiling
now adds the workspace's own bytes. Plan-reuse execute! stays 0 B.
Full suite: 32151 pass (was 32300), 20.5 min.

Cost of dropping the pool (workstation, taskset 4-7, medians; before ->
after):
  call floor, DefaultAllocator, ns (bytes after; before 0 B):
    Float64    8^3 336 -> 1023 (3.4 KB), 16^3 432 -> 2125 (6.9 KB),
               d4 outer 323 -> 754 (2.6 KB)
    ComplexF64 8^3 407 -> 1263 (4.1 KB), 16^3 940 -> 2917 (10.7 KB),
               d4 outer 545 -> 1046 (2.8 KB)
  @tensor n^3 matmul, DefaultAllocator, us:
    Float64    32: 1.26 -> 4.2, 64: 6.7 -> 16.1, 128: 46 -> 73,
               256: 382 -> 500 (0.94 MB/call)
    ComplexF64 32: 5.5 -> 11.6, 64: 27 -> 45, 128: 199 -> 270,
               256: 1420 -> 1600 (1.47 MB/call)
  @tensor n^3 matmul, Bumper default_buffer, us:
    Float64    32: 2.6 -> 2.0, 64: 10.0 -> 9.6, 128: 62 -> 61, 256: 456 -> 463
    ComplexF64 32: 6.9 -> 6.7, 64: 32 -> 32, 128: 227 -> 228, 256: 1600 -> 1600
  TTFX first call: Float64 3.6 -> 3.3 s, ComplexF64 2.2 -> 2.0 s.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A complex kernel keeps the k_block of real(T)'s default real kernel (B
sliver at half L1d) and derives m_block/n_block from its own packed bytes
at that depth, rounded down to MR/NR. The per-kernel model from chunk 11
let a complex B sliver halve k_block, which measured slower.

Same-node A/B over the planar, 1m and FMAddSub menu shapes (512^3,
2048^3, 12/24x2048x2048; Genoa, Rome, Icelake): hybrid/old 0.99-1.00
(worst 1.02), per-kernel model/old up to 1.10, worst at small M.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Blocking resolution (defaults, rounding, clamping, C panel, line-by-line
packing) and path selection move before the kernel barrier, computed from the
kernel type and tile (`resolve_blocking`, `select_path`). The barrier crosses
once with the kernel (`Val{Kern}()`, or a named kernel itself), the path
singleton, the transforms and one `RefValue` holding the request plus the
resolved blocking. `ContractPlan` gains a `path` field/type parameter;
`execute!` keeps its runtime short-circuits and then calls
`execute_path!(plan, ..., plan.path)` statically, so barrier #2 is gone.

The dot path's capacity check is evaluated at plan time (`dot_fits`) from
`workspace_sizes`, which now works on the kernel type and tile.
`pack_split` takes (side, lanes, format) instead of a `SliverSpec`.

Deleted: `_path_hint`, `_default_k_block`, `_continue`, `_execute_hinted!`,
`_hint_holds`, `_splits`, `_split_path`, `_execute_across_barrier!`,
`_execute_resolved!`, `_strip_storage`, `_with_storage`, `SlotCache`,
`barrier_slot!`/`barrier_slot_slow!`, the workspace's `slots` field,
`_dot_capacity_ok`, and the global `_DOT_MODE`/`_OUTER_MODE`/
`_UNPACKED_B_MODE` Refs (now an internal `path_modes::PathModes` keyword of
`plan_contract`). `Execute` is a callable struct (execute! + release!), and
plan construction ends with `req.f(plan)`. One keyword copy method
`ContractPlan(p; ...)` replaces the spelled-out copies in the panel path and
tests. Underscores dropped on touched names; barrier.jl -> paths.jl.

Measurements (workstation, cascadelake; noisy, medians of 4 interleaved runs
of per-call medians, indicative only), before -> after:
  execute! 64^3 Float64              6.8 us -> 6.8 us, 0 B
  execute! 2x2 Float64               127 ns -> 85-190 ns (bimodal), 0 B
  execute! 32^3 ComplexF64           5.4 us -> 5.3 us, 0 B
  execute! dot 1x256x64 Float64      1.77 us -> 1.74 us, 0 B
  execute! panel 32x600x32 F32/F64   30.5 us -> 30.3 us, 0 B
  plan_contract 64^3 Float64         1.94 us -> 2.06 us; 72240 -> 71952 B
  plan_contract 2x2 Float64          417 ns -> 405 ns; 2272 -> 1984 B
  tensorcontract! 2x2 Float64        486 ns -> 484 ns; 2096 -> 1808 B
  tensorcontract! 2x2 ComplexF64     687 ns -> 627 ns; 2432 -> 2144 B
TTFX (fresh process, median of 3): first Float64 plan+execute 3.44 ->
3.52 s, then ComplexF64 1.84 -> 1.93 s; first @tensor call 2.55 -> 2.53 s.
Codegen on a fixed 172-case workload (all paths, eltypes, conj, swap, split,
panel, named kernels, TensorOperations): path/nest/micro-tile/dot/outer
specializations identical (58/40/18/22/14/17/4), --trace-compile 715 -> 714
lines, except that k_length == 0 plans now compile their (never run) nest
path: +4 nests with four such cases included. Results of all 172 cases are
bitwise identical to before, with identical paths, blocking and splits.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
lkdvos and others added 20 commits October 6, 2026 15:18
`NestPath{UNPACKED_B, AFF, SPLIT}` becomes `NestPath{UNPACKED_B, SPLIT}`.
Sliver axes always take the runtime `axis_of` choice (the `Val`-typed
`axis_of` methods and `throw_irregular_ramp_descriptor` go, as do the
`Val(aff_...)` arguments of `micro_tiles!` and the `AFF` lookups in
`pack_block!`). The closed-form ramp descriptors stay: `nest!` decides
per group, once per call, with `affine_ramp` on the groups it walks (a
split group is not a ramp; under `PanelPath` the plan's C maps are the
panel's dense ones; under unpacked B, N's ramp needs B's and C's maps,
as before). The 512-entry `_NEST_PATHS`/`_PANEL_PATHS` tables,
`nest_path_from_flags` and `is_ramp_map` go; `nest_path(unpacked_b,
split_a, split_b, panel)` branches over the six valid combinations, one
literal singleton per arm. `benchmark/bench_ramp_flags.jl` goes, its
question answered; `submit_ab.sh` remains for A/B runs.

Slurm evidence (same-node A/B): round 1 (7183225-7) and the unpacked-B
codegen fix (ea97e6d, 7184013-5); round 2 (7184157-9, ec3d30a vs
3b2e48e on exp/runtime-axis-types): closed-form descriptors with runtime
axis types match the static flags within noise (+-4%), except F64
partial-ramp 1.07 and ccsd(t) 1.05 on Icelake.

Equivalence vs 470f977, 743 cases: all bitwise identical; paths
identical modulo the removed parameter (base had 4 distinct AFF
patterns across its nest paths).

Specialisations on the equivalence workload, before -> after: nest!,
micro_tiles!, b_sliver 50 -> 50, execute_path! 94 -> 94, build_plan
194 -> 194, pack_block! 70 -> 72, pack_sliver!/pack! 42 -> 84,
execute_micro_tile!/execute_tile! 43 -> 160, store_tile! 32 -> 116.
The workload's plan types each met few AFF patterns, so the nest-level
counts do not drop; the leaves grow because a ramp plan now compiles
the ScatterAxis arms too (4 row/column axis-type pairs per tile).
TTFX (fresh process, medians of 3, workstation): f64 64^3 3.30 -> 5.60 s,
c64 64^3 1.88 -> 2.66 s, @tensor 1.17 -> 2.19 s, using 0.46 -> 0.55 s.
Full test suite 22m12s -> 31m50s (32160 pass).

Per-call execute! (workstation, pinned, interleaved, medians of 3 runs
of 101-sample medians; indicative), new/base: 24^3 F64 0.99, C64 1.02;
64^3 F64 0.99, C64 1.00; 12x512x512 F64 0.98, C64 0.96; 512^3 F64 1.05;
split-A (intensli_7 16^6) 0.86 (noisy); panel (F32, Float64
accumulator, 64x700x40) 1.05; ccsd(t)-like permuted F64 1.01, C64 1.01.
0 allocations in all cases.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This reverts commit b77152b. With the ramp decision at runtime every
dense plan also compiles the scatter branches of the micro-tile, store
and pack code (execute_tile! 43 -> 160, store_tile! 32 -> 116
specialisations): TTFX +40-90% and the suite 22 -> 32 min, for no
runtime gain. The static flags pay for themselves in compile time.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
`planned` takes α and β as ordinary arguments (`nothing` = build the plan
only) and carries them in `PlanRequest` across the kernel barrier instead
of a continuation `f`; after the barrier `build_plan` returns the plan, or
runs `execute!` and `release!` and returns nothing (`execute_request`, two
methods on the field types). `Execute` goes. `plan_contract` passes
`nothing`; `contract!` and `TO.tensorcontract!` convert α/β to the compute
type and call `planned` directly, so `contract!` crosses one dynamic
dispatch instead of two. `planned`'s options become keywords with
`plan_contract`'s defaults (allocations and per-call floor unchanged).

Checks: the conjugated-output check lives only in `planned`; the aliasing
check moves there too, so `plan_contract`/`contract!` reject an output
sharing memory with an input (new lean test). TO's arg/dim checks stay in
the adapter. Renames: `_qs_labels` -> `contraction_labels`,
`_qs_prepare` -> `prepare_contraction`, `_qs_check_eligible` +
`_qs_throw` -> `check_strided`. The stale Zero()/One() comment now points
at `static_beta`.

Equivalence vs 65b3af0, 1134 cases (the existing 743: 731 plan/execute
and 12 tensorcontract!; plus 3 permuted/mixed/accumulator tensorcontract!
and 388 contract! calls): all bitwise identical, paths/blocking/splits identical.

Per-call (workstation, pinned, interleaved base/new processes, medians
of 10 runs of 101-sample medians; indicative), base -> new:
tensorcontract! 2x2 F64 454 -> 457 ns, C64 698 -> 695 ns (allocs
1856/2192 B both); contract! 2x2 F64 530 -> 455 ns, 2064 -> 1856 B;
contract! 64^3 F64 25.3 -> 24.5 us, 72032 -> 71824 B; execute! on a
reused plan 76 -> 76 ns (2x2), 6.74 -> 6.75 us (64^3), 0 B. 64^3 calls
are bimodal (~12.5 / ~19-25 us, fresh-workspace memory); one case per
process, medians of 6: tensorcontract! 64^3 15.8 -> 12.6 us, contract!
64^3 21.3 -> 15.1 us, same two modes in both, 71824 B both.

TTFX (fresh process, 3 interleaved runs): first plan+execute F64 64^3
3.21 -> 3.25 s, C64 1.88 -> 1.86 s, first @tensor 1.19 -> 1.13 s, first
contract! (F32) 1.14 -> 1.07 s. build_plan specialisations 8 -> 8 in
that session. On the equivalence workload build_plan 200 -> 304: a
`contract!` call no longer shares its plan type's build_plan with
`plan_contract` (as `tensorcontract!` already did not); everything below
build_plan (execute_path!, nest!, packers, tiles) unchanged.

Full suite: 32163 pass (19m14s); Runic clean.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
ParallelTestRunner runs every test/**/test_*.jl in its own module on a pool
of worker processes; arguments select files by name prefix. Each file
includes test/helpers.jl or its folder's helpers.jl, so it also runs on its
own, and no file depends on another. Test files follow src/: the per-call
overhead file is split into labels, axis_group, tiles, nest, execute and
plan; label/K-order tests go to test_labels.jl, blocking tests to
test_blocking.jl, run-length demotion to test_kernel_selection.jl, line
packing to test_line_packing.jl, the C panel to test_c_panel.jl. Every
@test is kept unchanged.

forced_isa_runner.jl now calls Pkg.test with QS_FAKE_ISA set, and
runtests.jl applies the fake target in each worker.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The audit's trims 1-9:
- dot and outer paths: compared against the reference only, not also
  against the nest (both checks were tolerance checks).
- The ScalarKernel-vs-SIMDKernel execute! testset is gone: the randomized
  nest test runs the same four Float64 kernels against a dense reference.
- Label/K-order correctness sweeps (ccsd_t, scrambled K) run Float64 and
  ComplexF64; the adapter-path label test runs Float64.
- The randomized nest test keeps random conj flags, blockings, alpha/beta
  and NaN C but no longer randomizes StridedView ops; the XOR rule stays
  pinned by its own testsets.
- No mk_e2e runs in the planar, 1m and FMAddSub kernel files; the
  mixed-kernel end-to-end loop keeps ComplexReal ComplexF64 (:always,
  :never) and RealComplex ComplexF32, the only nest runs of their kind.
- TensorOperations: @tensor/ncon and flags x op on ComplexF64 only; the
  StridedNative comparison uses the diagonal conj pairs for complex and
  (false, false) for real eltypes.
- Merged duplicates: the manual-pipeline bounds/shortcut testset into
  mk_contract_full (destination BoundsError before any write; alpha =
  beta = 0 over a NaN C); the second randomized affine_ramp check; the two
  AxisGroup-from-labels testsets (20 trials, ranks 1-4); the 1m/FMAddSub
  kernel_from_shape loops into kernel_selection, their blocking ratios
  into test_blocking; FMAddSub's interleaved-A packing (test_pack_* cover
  it); the default-kernel testset into the plain-GEMM demotion check
  (keeping contract! == execute!); the conjugated-C nest testset (a
  transpose op on a complex C is now checked in test_plan); the
  scattered-default allocation check into test_execute, which now covers
  Float32 there too.

Coverage, per-testset over every file (before / after): 1759 / 1759 src
lines, none lost. Every eltype x kernel family pair (14) and every
individual NestPath flag value per family, and per eltype and family,
still reaches nest!; 11 of 86 coarse (eltype, family, flag tuple)
combinations go, each recombining flag values covered elsewhere. Every
shipped menu kernel compiled before is still compiled (add_tile,
execute_tile!, store_tile!). Lost pack! shapes: 1e A at MR = 12
(ComplexF64) and 24 (ComplexF32), from the dropped 1m end-to-end runs,
plus test-only shapes.

Pkg.test: 3m38s -> 2m53s wall, 32164 -> 30888 tests; with
--check-bounds=auto 2m25s -> 2m07s. Slowest files (s, before -> after):
test_tensoroperations 183 -> 146, test_labels 173 -> 117, test_unpackedb
157 -> 148, test_line_packing 139 -> 129, test_onem_kernel 124 -> 115,
test_dot 118 -> 60.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The 1m end-to-end runs dropped in 6b2c50d were the only tests compiling
the 1e A packing at MR = 12 (ComplexF64) and 24 (ComplexF32); the fast-path
test now covers every 1m menu MR instead of the menu extremes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Internal constants drop the leading underscore, like the functions in
earlier chunks; _QS_ELTYPES becomes SUPPORTED_ELTYPES.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…lags

Delete bench_complex_blocking.jl and bench_ramp_flags.jl (answered). The
probes become the standard checks: probe_ttfx.jl times load, the first
Float64/ComplexF64 plan+execute!, @tensor and contract!; probe_call_floor.jl
adds contract! and reused-plan execute! (2x2, 64^3); new
probe_specialisations.jl counts specialisations of the hot functions after
a fixed workload. All take --label; submit_ab.sh copies a script a revision
lacks from the submitting checkout. Every bench_*.jl takes --smoke (helpers
in harness.jl); blocking-model-only helpers move into that script. Fix the
blocking model summary passing a dtype to named_rows. Add benchmark/README.md.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
plot_bench_to_suite.jl now draws two figures with one panel per dtype:
throughput against arithmetic intensity (colour = total work, filled =
QuasiStrided, open = StridedBLAS, the two joined per case) and the
QuasiStrided/StridedBLAS time ratio against total work (colour = intensity,
marker = category), with the geomean and faster count in the panel title
and the three most extreme cases at each end numbered and named below.
--dtypes, --categories and --tags filter the rows; the filters name the
output files. The per-group figures and the synthetic violins are gone.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
plot_bench_to_suite.jl draws a single figure: one column per dtype, three
rows on a shared arithmetic-intensity axis (QuasiStrided GFLOP/s,
StridedBLAS GFLOP/s on the same scale, and the time ratio with the geomean
and faster count). Colour = total work on a trimmed batlow map with one
colorbar, marker = category, semi-transparent markers; the pairing segments
and the extreme-case labels are gone. Log axes get plain tick labels spaced
to their span. Output: bench_to_suite_overview[_<filters>].png.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
One row per dtype, columns QuasiStrided GFLOP/s | StridedBLAS GFLOP/s |
time ratio on the linked arithmetic-intensity axis; the two throughput
panels of a row share y-limits. The geomean and faster count move into a
boxed label in the ratio panel's corner. Without --dtypes the plot shows
whichever of Float64 and ComplexF64 the CSV has, else every dtype in it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Its decisions are summarised in the pull request; open ideas are issues
#13-#15, #17-#21.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The geomean box sat over the lowest ratios and could hide an outlier;
the y-range now leaves an empty band below the data for it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@lkdvos

lkdvos commented Oct 7, 2026

Copy link
Copy Markdown
Member Author

Upstream suite after the cleanup

bench_to_suite.jl, job 7188915 (Icelake, single core, commit 5189efa), Float64 and ComplexF64, :contract and :network; no mismatches or backend errors, canary spread 1.5%.

geomean QS/BLAS time QuasiStrided faster
Float64 0.83 214 of 300
ComplexF64 0.78 236 of 293
group Float64 ComplexF64
network mps / ctmrg / trg 0.79 / 0.79 / 0.59 0.67 / 0.72 / 0.53
synthetic, permuted or scrambled 0.63–0.85 0.67–0.84
tccg, not BLAS-equivalent 0.58 0.65
synthetic gemm_ready 1.22 0.93
tccg, BLAS-equivalent (ccsd_1) 2.49 1.32
batched 3.27 1.57

The network cases, which an earlier complex-only run (older code) showed at about 1.1–1.2×, are now clearly faster. The remaining losses are per-call overhead on batched and tiny BLAS-shaped contractions: #23.

🤖 Generated with Claude Code

@lkdvos
lkdvos merged commit ad94b73 into main Oct 7, 2026
@lkdvos
lkdvos deleted the cleanup branch October 7, 2026 17:21
lkdvos added a commit that referenced this pull request Oct 7, 2026
At 16 ymm registers LLVM's haswell/znver2 scheduler hoists every B broadcast
of a planar or fmaddsub K step to the top of the loop body, holds them all
live and spills them. `kstep_fence()`, an empty `asm sideeffect` with a
memory clobber (no instruction), after each column group pins the loads next
to their FMAs; the FMAs stay free to move. On AVX2, fmaddsub also runs
OpenBLAS's Zen zgemm order: the swap(a)*bi pass, a fence, then the a*br pass
on A loaded again, so a and swap(a) are never live together.

Only 256-bit kernels (`W * sizeof(R) == 32`) take the fenced body;
every other shape runs the unchanged generator. code_llvm/code_native of
`add_tile` are identical to ad94b73 for all 57 non-256-bit complex and
mixed-domain menu shapes (B packed and B in place) under -C icelake-server,
znver4, znver2 and haswell.

Hot-loop stack references per K step, Julia 1.12.7, -C znver2 (haswell
identical), ad94b73 -> this commit, B packed / B read in place (the #22
unpacked-B codegen does not change the picture):
  planar   4x5/W4, 8x5/W8     4 / 4   ->  0 / 0
  fmaddsub 4x6/W4, 8x6/W8    24 / 24  ->  0 / 0
  fmaddsub 4x5/W4, 8x5/W8     0 / 0   ->  0 / 0
  planar   4x6/W4, 8x6/W8    14 / 12  ->  0 / 13  (not in any menu)
The mixed-domain kernels run the real SIMDKernel step and do not spill.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
lkdvos added a commit that referenced this pull request Oct 7, 2026
… vector store

On AVX2, `select_shape` hands a complex contraction to FMAddSub 4x6/W4
(ComplexF64) or 8x6/W8 (ComplexF32) when C's unit run along M fills whole
`MR`-row slivers (`run == m_length || run % MR == 0`); otherwise it keeps
planar 4x5/8x5. FMAddSub's single accumulator plane holds 12 accumulators +
2 A + 1 broadcast, 15 of 16 registers, and stays spill-free with B packed or
in place. Planar NR = 6 fills all 16 and spills 13 per K step with B in
place, so it is not a candidate. Where C's rows defeat the vector store,
fmaddsub's lane-pair scalar store costs about 1.3x planar's per element,
and with a short K the store dominates (ccsd_t_2/3/4, ao2mo_2/3), hence the
run condition. `plan_contract` now computes the run for complex `T` too.

Plans over the 490 suite :contract cases with uniform legs, same-eltype and
both mixed-domain orientations, icelake-server/znver4/znver2 profiles,
ad94b73 vs this branch: identical everywhere except 419 same-eltype plans on
znver2, which change kernel (planar -> fmaddsub) and the blocking derived
from it; the execution path is unchanged.

Rome, same-node A/B of the same change on 4a77066, before the cleanup (#22)
(jobs 7143376/7143377, two rounds each; GF/s of the default kernel, main
planar 4x5 -> fmaddsub 4x6, OpenBLAS in brackets):
  ComplexF64 64^3 35.9 -> 40.6 [42.3], 256^3 42.1 -> 44.6 [49.3],
             512^3 41.2 -> 48.2 [50.3], 2048^3 40.5 -> 45.1 [51.2],
             smallMN 29.3 -> 37.0 [30.2]; time geomean 0.87, no case slower
  ComplexF32 64^3 61.1 -> 76.7 [72.5], 256^3 80.7 -> 90.4 [95.2],
             2048^3 81.5 -> 99.3 [102.8], smallN 44.0 -> 73.0 [53.1];
             time geomean 0.79, no case slower
Suite (complex :contract and :network): QuasiStrided time geomean 0.78
(ComplexF32) / 0.83 (ComplexF64) of main. Consistently slower: the
dim32/96_1_3_1 contract_scrambled cases (1.23-1.47), and in ComplexF64
intensli_7_dim24 and dim8-16 0_2_3 cases (1.06-1.14). A diagnostic run
traced the scrambled cases to the kernel method (fmaddsub), not the fence.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
lkdvos added a commit that referenced this pull request Oct 8, 2026
At 16 ymm registers LLVM's haswell/znver2 scheduler hoists every B broadcast
of a planar or fmaddsub K step to the top of the loop body, holds them all
live and spills them. `memory_fence()`, an empty `asm sideeffect` with a
memory clobber (no instruction), after each column group pins the loads next
to their FMAs; the FMAs stay free to move. On AVX2, fmaddsub also runs
OpenBLAS's Zen zgemm order: the swap(a)*bi pass, a fence, then the a*br pass
on A loaded again, so a and swap(a) are never live together.

Only 256-bit kernels (`W * sizeof(R) == 32`) take the fenced order; every
other shape gets the same statements in the same order as before. code_llvm/code_native of
`add_tile` are identical to ad94b73 for all 57 non-256-bit complex and
mixed-domain menu shapes (B packed and B in place) under -C icelake-server,
znver4, znver2 and haswell.

Hot-loop stack references per K step, Julia 1.12.7, -C znver2 (haswell
identical), ad94b73 -> this commit, B packed / B read in place (the #22
unpacked-B codegen does not change the picture):
  planar   4x5/W4, 8x5/W8     4 / 4   ->  0 / 0
  fmaddsub 4x6/W4, 8x6/W8    24 / 24  ->  0 / 0
  fmaddsub 4x5/W4, 8x5/W8     0 / 0   ->  0 / 0
  planar   4x6/W4, 8x6/W8    14 / 12  ->  0 / 13  (not in any menu)
The mixed-domain kernels run the real SIMDKernel step and do not spill.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
lkdvos added a commit that referenced this pull request Oct 8, 2026
… vector store

On AVX2, `select_shape` hands a complex contraction to FMAddSub 4x6/W4
(ComplexF64) or 8x6/W8 (ComplexF32) when C's unit run along M fills whole
`MR`-row slivers (`run == m_length || run % MR == 0`); otherwise it keeps
planar 4x5/8x5. FMAddSub's single accumulator plane holds 12 accumulators +
2 A + 1 broadcast, 15 of 16 registers, and stays spill-free with B packed or
in place. Planar NR = 6 fills all 16 and spills 13 per K step with B in
place, so it is not a candidate. Where C's rows defeat the vector store,
fmaddsub's lane-pair scalar store costs about 1.3x planar's per element,
and with a short K the store dominates (ccsd_t_2/3/4, ao2mo_2/3), hence the
run condition. `plan_contract` now computes the run for complex `T` too.

The FMAddSub NR = 5 entries (4,5,4)/(8,5,8) leave the menus: nothing selects
them, and the fenced NR = 6 tiles are spill-free. One `shape_override` row
now gives both lane-pair kernels (1m, FMAddSub) their AVX2 shape; for
OneMKernel ComplexF32 that is (8,6,8) where the fit used to hand it the
AVX-512 menu head (24,8,16).

Plans over the 490 suite :contract cases with uniform legs, same-eltype and
both mixed-domain orientations, icelake-server/znver4/znver2 profiles,
ad94b73 vs this branch: identical everywhere except 419 same-eltype plans on
znver2, which change kernel (planar -> fmaddsub) and the blocking derived
from it; the execution path is unchanged.

Rome, same-node A/B of the same change on 4a77066, before the cleanup (#22)
(jobs 7143376/7143377, two rounds each; GF/s of the default kernel, main
planar 4x5 -> fmaddsub 4x6, OpenBLAS in brackets):
  ComplexF64 64^3 35.9 -> 40.6 [42.3], 256^3 42.1 -> 44.6 [49.3],
             512^3 41.2 -> 48.2 [50.3], 2048^3 40.5 -> 45.1 [51.2],
             smallMN 29.3 -> 37.0 [30.2]; time geomean 0.87, no case slower
  ComplexF32 64^3 61.1 -> 76.7 [72.5], 256^3 80.7 -> 90.4 [95.2],
             2048^3 81.5 -> 99.3 [102.8], smallN 44.0 -> 73.0 [53.1];
             time geomean 0.79, no case slower
Suite (complex :contract and :network): QuasiStrided time geomean 0.78
(ComplexF32) / 0.83 (ComplexF64) of main. Consistently slower: the
dim32/96_1_3_1 contract_scrambled cases (1.23-1.47), and in ComplexF64
intensli_7_dim24 and dim8-16 0_2_3 cases (1.06-1.14). A diagnostic run
traced the scrambled cases to the kernel method (fmaddsub), not the fence.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
lkdvos added a commit that referenced this pull request Oct 8, 2026
…k/store (#16)

* Fence the AVX2 complex K steps so their broadcasts stop spilling

At 16 ymm registers LLVM's haswell/znver2 scheduler hoists every B broadcast
of a planar or fmaddsub K step to the top of the loop body, holds them all
live and spills them. `memory_fence()`, an empty `asm sideeffect` with a
memory clobber (no instruction), after each column group pins the loads next
to their FMAs; the FMAs stay free to move. On AVX2, fmaddsub also runs
OpenBLAS's Zen zgemm order: the swap(a)*bi pass, a fence, then the a*br pass
on A loaded again, so a and swap(a) are never live together.

Only 256-bit kernels (`W * sizeof(R) == 32`) take the fenced order; every
other shape gets the same statements in the same order as before. code_llvm/code_native of
`add_tile` are identical to ad94b73 for all 57 non-256-bit complex and
mixed-domain menu shapes (B packed and B in place) under -C icelake-server,
znver4, znver2 and haswell.

Hot-loop stack references per K step, Julia 1.12.7, -C znver2 (haswell
identical), ad94b73 -> this commit, B packed / B read in place (the #22
unpacked-B codegen does not change the picture):
  planar   4x5/W4, 8x5/W8     4 / 4   ->  0 / 0
  fmaddsub 4x6/W4, 8x6/W8    24 / 24  ->  0 / 0
  fmaddsub 4x5/W4, 8x5/W8     0 / 0   ->  0 / 0
  planar   4x6/W4, 8x6/W8    14 / 12  ->  0 / 13  (not in any menu)
The mixed-domain kernels run the real SIMDKernel step and do not spill.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Open the complex pack and store fast paths to AVX2

`complex_fastpath_isa_eligible` was AVX-512 only, so on AVX2 every complex
sliver packed through the scalar loop and every complex tile (planar,
fmaddsub and the mixed-domain kernels, which share those stores) stored
element by element. The AVX2 codegen is the AVX-512 path at half width.
NEON and unknown targets keep the scalar loops.

Measured on the pre-refactor branch (Cascade Lake, QS_FAKE_ISA=avx2
-C znver2, gate off -> on, GF/s): kernel-only planar 8x6/W8 85.4 -> 108.5;
engine ComplexF32 2000^3 94.2 -> 105.1, shallowK 47.0 -> 85.5, smallN
47.8 -> 70.4.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* AVX2 complex default: fmaddsub 4x6/W4, 8x6/W8 where C's rows take its vector store

On AVX2, `select_shape` hands a complex contraction to FMAddSub 4x6/W4
(ComplexF64) or 8x6/W8 (ComplexF32) when C's unit run along M fills whole
`MR`-row slivers (`run == m_length || run % MR == 0`); otherwise it keeps
planar 4x5/8x5. FMAddSub's single accumulator plane holds 12 accumulators +
2 A + 1 broadcast, 15 of 16 registers, and stays spill-free with B packed or
in place. Planar NR = 6 fills all 16 and spills 13 per K step with B in
place, so it is not a candidate. Where C's rows defeat the vector store,
fmaddsub's lane-pair scalar store costs about 1.3x planar's per element,
and with a short K the store dominates (ccsd_t_2/3/4, ao2mo_2/3), hence the
run condition. `plan_contract` now computes the run for complex `T` too.

The FMAddSub NR = 5 entries (4,5,4)/(8,5,8) leave the menus: nothing selects
them, and the fenced NR = 6 tiles are spill-free. One `shape_override` row
now gives both lane-pair kernels (1m, FMAddSub) their AVX2 shape; for
OneMKernel ComplexF32 that is (8,6,8) where the fit used to hand it the
AVX-512 menu head (24,8,16).

Plans over the 490 suite :contract cases with uniform legs, same-eltype and
both mixed-domain orientations, icelake-server/znver4/znver2 profiles,
ad94b73 vs this branch: identical everywhere except 419 same-eltype plans on
znver2, which change kernel (planar -> fmaddsub) and the blocking derived
from it; the execution path is unchanged.

Rome, same-node A/B of the same change on 4a77066, before the cleanup (#22)
(jobs 7143376/7143377, two rounds each; GF/s of the default kernel, main
planar 4x5 -> fmaddsub 4x6, OpenBLAS in brackets):
  ComplexF64 64^3 35.9 -> 40.6 [42.3], 256^3 42.1 -> 44.6 [49.3],
             512^3 41.2 -> 48.2 [50.3], 2048^3 40.5 -> 45.1 [51.2],
             smallMN 29.3 -> 37.0 [30.2]; time geomean 0.87, no case slower
  ComplexF32 64^3 61.1 -> 76.7 [72.5], 256^3 80.7 -> 90.4 [95.2],
             2048^3 81.5 -> 99.3 [102.8], smallN 44.0 -> 73.0 [53.1];
             time geomean 0.79, no case slower
Suite (complex :contract and :network): QuasiStrided time geomean 0.78
(ComplexF32) / 0.83 (ComplexF64) of main. Consistently slower: the
dim32/96_1_3_1 contract_scrambled cases (1.23-1.47), and in ComplexF64
intensli_7_dim24 and dim8-16 0_2_3 cases (1.06-1.14). A diagnostic run
traced the scrambled cases to the kernel method (fmaddsub), not the fence.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant