Repository navigation
Conversation
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Decision D1: names follow `<m|n|k>_<level>_<what>` codebase-wide. Public renames: - `Blocking` fields `mc, kc, nc` -> `m_block, k_block, n_block` - `plan_contract` keywords `mc/kc/nc` -> `m_block/k_block/n_block` - `mr(kernel)`, `nr(kernel)` -> `tile_size(kernel) -> (m_tile, n_tile)` - `packed_a_per_k`, `packed_b_per_k` -> `sliver_widths(kernel)` (`tile_size` and `sliver_widths` are now in the `public` list) `using LinearAlgebra` is dropped from src: it used no LinearAlgebra names (`adjoint`/`transpose` are Base's). Full test suite: 34542/34542 pass. TTFX unchanged (Float64 first call 3.3-3.5 s before, 3.2-3.7 s after). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
`sliver_widths` becomes `sliver_width`, matching `size`; both accessors gain an index method, replacing the `tile_size(k)[i]` spelling. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- The ISA is the CPUID probe's; the LLVM CPU-name table is gone (it could only narrow the probe, e.g. capping AVX-512 Zen 4/5 at AVX2 on Julia 1.10). - `CacheLevel` drops the unused `ways`; `TargetProfile` drops `arch` and derives `vector_bytes`/`nregisters` from `isa`, rejecting unknown ISAs. - `core_bytes(profile, level)` replaces three copies of the per-core cache share in blocking.jl, pack_split.jl and labels.jl. - No leading underscores in target.jl; `_TARGET` is `TARGET`. - Tests: synthetic profiles take only the ISA, so the sweeps over impossible (isa, vector_bytes, nregisters) combinations go (34542 -> 33444 tests); the forced-ISA runner takes QS_FAKE_ISA without QS_FAKE_VB/QS_FAKE_NREG. Full suite passes (15 min). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- `pair_group.jl` becomes the `AxisGroup(labels, (ind1, v1), (ind2, v2))` constructor. - `normalize_group` is gone and `offsets` moves to the test helpers (both were test-only). - `map_ramp_step` moves from dot.jl; `affine_ramp` is built on it. - `Base.setindex`, `prod` and `Base.Checked` replace hand-written versions. - `BlockDescriptor`/`describe_block` move to tiles.jl; `describe_block` checks `count <= length(buffer) - first`, so `first + count` cannot overflow past the check. - `fill_offsets!` drops its alias check. - No leading underscores in axis_group.jl. Full suite passes (30989 tests, 15 min; the drop is the removed normalize_group/offsets/alias tests). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- AffineAxis <: AbstractVector{Int} (one-based getindex, closed-form
checked extrema); the scatter axis is a plain view of the offset buffer.
ScatterAxis, PtrScatterAxis, axis_from_descriptor and the Axis union go.
- Tile{S, R, C} replaces QSTile/SourceTile/DestinationTile; size and
@propagate_inbounds getindex/setindex! replace nrows/ncols/tile_offset/
tile_load/tile_store!, with @inbounds at the former call sites.
- extrema replaces axis_offset_range; is_unit_stride (with an
AbstractVector fallback) replaces _unit_stride_rows.
- Tile coordinates and axis_offset stay zero-based for now (temporary shims).
- The nest passes axes to their consumers through _with_axis, one branch
per axis type: a Union holding a view is heap-boxed (432 B to 1.6 KB per
call on a contraction with scattered M, N and K). A zero-allocation test
for that contraction is added.
Tests: 30930 pass (30989 before; removed ScatterAxis/PtrScatterAxis/
axis_from_descriptor tests).
TTFX probe, Float64 / ComplexF64 first call, interleaved runs:
before 3.1-3.3 s / 2.1-2.3 s, after 3.4-3.7 s / 2.3-2.4 s.
Reused-plan execute! medians (us, 3 interleaved runs, before -> after):
matmul 256^3 Float64 404 -> 395, ComplexF64 1537 -> 1442;
scattered-M 12x20x256x256 Float64 432 -> 431, ComplexF64 1460 -> 1456;
scattered-M/N/K small Float64 17.2 -> 17.6, ComplexF64 40.4 -> 39.1.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Tile indexing, AffineAxis/view axes, packer lane and K-step loops,
microkernel K steps, columns and row lanes (generated code included) and
UnpackedBView columns/K steps are one-based; the two zero-based shims go.
- packed_{a,b}_offset and packed_{a,b}_plane_offset take one-based
coordinates and still return zero-based offsets: the panel layout is
unchanged. Addresses, block starts and AxisGroup coordinates stay
zero-based, as does the macro-block coordinate of the line-by-line packer.
Tests: 30930 pass.
TTFX probe, Float64 / ComplexF64 first call, interleaved with the chunk-2
commit: before 3.1-3.2 s / 2.1-2.2 s, after 3.3-3.4 s / 2.2-2.4 s.
Reused-plan execute! medians (us, 3 interleaved runs, before -> after):
matmul 256^3 Float64 411 -> 400, ComplexF64 1564 -> 1532;
scattered-M 12x20x256x256 Float64 448 -> 431, ComplexF64 1503 -> 1499;
scattered-M/N/K small Float64 17.9 -> 17.6, ComplexF64 42.0 -> 41.4.
bench_kernels.jl kernel arm, Float64, median of 5 interleaved runs (GF/s):
16x6/W8 85.8 -> 84.7, 32x6/W8 107.0 -> 106.9.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A `view` scatter axis is not `isbits`, so the nest's
`Union{AffineAxis, view}` was heap-boxed, which `_with_axis` worked around
with about ten nested closures. `ScatterAxis` is again a borrowed pointer,
now an `AbstractVector{Int}` indexed one-based, so `_axis_of` returns an
unboxed union and the nest reads flat. The oracle and `_scale_all_of_C!`
now `GC.@preserve` the workspace they borrow from.
Full suite passes (30930 tests). Reused-plan execute! medians, closures ->
pointer, us: matmul 256^3 F64 410 -> 410, C64 1578 -> 1559; scattered M F64
432 -> 441, C64 1545 -> 1531; scattered M/N/K F64 18.4 -> 18.1, C64 41.9 ->
42.0. Zero allocations per call in all six, under --check-bounds=yes. First
call F64/C64: 3.5-3.8 s / 2.4-2.8 s -> 3.2 s / 2.1-2.2 s (back to the
pre-chunk-3 3.1-3.3 s / 2.1-2.3 s).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Panels are only `PackedPanel`: the `AbstractVector` methods of `panel_vload`/`panel_load`/`panel_store!` are gone. The checked entry points `pack_a!`, `pack_b!` and `accumulate` take a `DenseVector` and borrow it as a `PackedPanel` under `GC.@preserve`; `execute_tile!` and the oracle reach vectors through `accumulate`. The kernels' `accumulate` methods constrain the A panel to `PackedPanel` (B may be an `UnpackedBView`), which keeps the vector method unambiguous. - One `Descriptor(Val(MR), Val(NR), T, a_format = RealFormat(), b_format = RealFormat())` constructor and a `RealDescriptor` alias replace `KernelDescriptor`/`ComplexKernelDescriptor`; `check_descriptor` names `Descriptor` in its messages. - `packed_a_offset(d, i, p, plane = 0)` / `packed_b_offset(d, j, p, plane = 0)` replace the plane and real-only offset functions; the layout is unchanged. The packers' `(plane, i, p)` closures adapt the order. - transposed.jl is merged into the end of panel.jl. - Tests: view destinations dropped (no longer supported), the unsafe packer test borrows its buffers, the "complex kernel has no single-plane offset" MethodError test is replaced by a plane-offset check. Measured against 81cfaa3 on a workstation: - Tests: 30931 pass (baseline 30930). - TTFX (load; first Float64, ComplexF64, K=1 outer call), 2 runs each, interleaved: before 0.5/2.9/2.0/0.6 s and 0.5/3.0/2.0/0.6 s, after 0.5/3.0/2.1/0.6 s and 0.5/2.9/2.0/0.6 s. - Reused-plan execute! medians over 5 interleaved processes (before -> after, us): matmul256 Float64 393.5 -> 391.8, ComplexF64 1472.7 -> 1487.2; scatterM Float64 414.9 -> 414.5, ComplexF64 1445.8 -> 1460.7; scatterMN Float64 17.5 -> 16.7, ComplexF64 39.7 -> 39.2 (minima equal). - 0 bytes allocated per call in every case under --check-bounds=yes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A and B are packed by one code path; the kernel's descriptor decides each
operand's packing through a `SliverSpec{I, L, F, T}` (side, lanes, format,
scalar type), from `sliver_spec(kernel, i)`.
- `pack!(panel, tile, spec, transform)` / `unsafe_pack!` replace
pack_a!/pack_b!/unsafe_pack_a!/unsafe_pack_b! and the DescriptorKernel
forwarding methods. Lanes run along `tile.rows`, K steps along `tile.cols`;
B is packed as `transpose` of its tile (the nest and oracle pass B's axes
swapped). `Base.transpose(::Tile)` added.
- `panel_offset(spec, t, p, plane)` is the one offset formula;
`packed_a_offset`/`packed_b_offset`/`packed_*_length` are defined on it
(layout unchanged).
- Fast paths are symmetric: real B with unit-stride columns now takes the
`Vec{L}` load/convert/store path.
- `emit!` takes `(re, im, neg_im)`; the value and padding callers supply them
(padding stays literal +0.0). The block packer takes the spec.
- One complex contiguous K loop with a per-format `store_step!` (LLVM IR of
the shuffles/loads/stores identical to before for all formats x transforms).
- Dead `PackedPanel` eligibility conditions dropped; eltype error built in an
`@noinline` helper; `is_unit_stride`, `dense_lanes`, `lane_convertible`,
`complex_fastpath_isa_eligible` collected in pack_contiguous.jl; no leading
underscores in pack.jl / pack_contiguous.jl.
Tests: 30931 -> 31537. Real A/B halves merged into sweeps over specs, plus
"B packs as A packs its transposed tile" and a real-B fast-path sweep
(L = 6); removed the duplicate ISA-gate testset (planar store), the
PackedPanel-gate eligibility checks, the fmaddsub Vector-vs-panel duplicate
and the unsafe_pack_b! duplicates.
TTFX (base / new, 2 interleaved runs): first call Float64 4.5, 3.4 / 4.2,
4.3 s; ComplexF64 3.1, 3.0 / 2.7, 2.7 s; load 0.5-0.6 s both.
A/B, reused-plan execute! medians (5 rounds, us, base -> new): matmul256
F64 513 -> 521, scatterM F64 600 -> 568, scatterMN F64 31.6 -> 32.0,
matmul256 C64 2005 -> 1981, scatterM C64 2011 -> 1963, scatterMN C64
62.1 -> 61.7. Row-major B (7 rounds): F64 571 -> 482, F32 282 -> 265
(5-round run: 605 -> 566, 232 -> 269; workstation noise is +-30%).
Scattered K (7 rounds): F64 582 -> 590, C64 2069 -> 1717. Isolated real B
sliver pack, fast vs scalar, kc = 256: 2-3x faster for every L in
4, 6, 8, 12, 16 (Float32 and Float64), so no power-of-two restriction.
Per-call allocations 0 for all cases under --check-bounds=yes.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- `Base.accumulate` -> `add_tile`, a function of our own, in every kernel,
the `DenseVector` method, callers and tests.
- `KernelMethod` (was `ComplexMethod`) singletons declare their packing
formats once with `pack_formats`; `a_reals`/`b_reals` go, blocking uses
`reals_per_element` of the formats. A test checks every kernel's descriptor
agrees with `pack_formats(complex_method(k))`.
- `DescriptorKernel` -> `Microkernel`; leading underscores dropped from the
names in interface.jl and scalar.jl.
- One `@inline execute_tile!` and one `@inline pack!`, whose storage check is
`@boundscheck`; every other validation stays unconditional. The nest calls
`@inbounds _pack_sliver!`/`@inbounds _execute_micro_tile!` (both
`@propagate_inbounds`, replacing `unsafe_execute_micro_tile!` and the
function-valued `pack!` argument); the oracle calls them checked. LLVM of
`_execute_nest!`, `_micro_tiles_packed_b!` and `_micro_tiles_unpacked_b!`
(Float64/ComplexF64, affine and scattered) has the same line count as before
and no per-tile storage check; under `--check-bounds=yes` the per-tile
checks now run, so the hoisting test counts them.
- `ScalarKernel` accumulates into an `NTuple{MR*NR,T}` (column-major) and its
store branches on `beta` once per tile; it now allocates nothing.
- Contract moved from the header comment into docstrings (`Microkernel`,
`KernelMethod`, `zero_accumulator`, `add_tile`, `store_tile!`,
`execute_tile!`, `pack!`).
Tests: 31548 pass (was 31537).
TTFX (2 runs each, interleaved), before and after alike: load 0.5 s; first
call Float64 3.2-3.3 s, ComplexF64 2.3 s; K=1 outer 0.6 s.
A/B medians of 5 interleaved rounds, after/before: matmul256 F64 0.997,
C64 0.991; scatterM 1.017/0.994; scatterMN 1.003/1.001; rowmajB256 F64
1.009, F32 0.996; scatterK F64 0.996, C64 1.009. Allocations 0 per call
under --check-bounds=yes.
ScalarKernel 4x4 nest, 20x30x25 output, K = 16: Float64 155 -> 50 us,
Float32 159 -> 46-55 us; allocations per call 218400 / 151200 bytes -> 0.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
All vector kernel structs move to src/microkernels/types.jl under an
abstract VectorKernel{MR, NR, T, W}, with three accumulator layouts
(RealLayout, SplitLayout, LanePairLayout) and the generator-time pieces
of the stores (index, lane value, full-block store, store prelude), so
they dispatch instead of `<:` tests. One generic add_tile loops
accumulate_step over k_steps (1m: 2k); 1m and the mixed-domain kernels
forward their step to a computed inner(kernel) (no inner field, no KI
parameter; the call folds to a singleton). One generic store_tile!,
vector store and scalar store parameterised by the layout; @inline per
layout as before (real/split yes, lane pairs no). One generic
constructor with the complex-T check and vector-shape check, generic
lanewidth and zero_accumulator. 1m gets the lane-pair vector store.
Renames: b_scalar/b_complex (UnpackedBView hooks), accumulate_step,
vector_store!/scalar_store!/vector_store_eligible, split_index,
split_store_block!/lanepair_store_block!, fmaddsub/swap_pairs/addsub.
Codegen (AVX-512, default shapes): K-loop instruction counts identical
for all six kernels; store_tile!/vector/scalar store instruction counts
identical (SIMD 2015/1953/1317, Planar 1511/1494/907, FMAddSub
101/3619/2423, ComplexReal 101/3612/2417, RealComplex 1976/1900/1202);
1m now has the 3612-instruction lane-pair vector store.
add_tile is now @inline for every kernel (before: only SIMD and mixed).
A/B vs 95728d0, medians of 5 interleaved runs: small-K (M = N = 200,
K = 1..8) ComplexF64 -23..-35%, Float64 -3..+2%; 256^3 named kernels:
SIMD -1%, Planar -2.5%, 1m -11%, FMAddSub -7%, ComplexReal +3%,
RealComplex 0%; earlier matmul/scattered/row-major/scattered-K cases
-4..+2%. Allocations 0 under --check-bounds=yes.
TTFX: Float64 first call 3.1 s (before 3.1/3.4 s), ComplexF64 2.1 s
(before 2.1/2.2 s). Tests: 31967 pass (+419 for the 1m vector store),
suite ~18 min.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
vector_store! generates its block loop once per beta case (iszero, isone, general) and branches once per tile; the full-block expressions are the specialised ones (real: alpha*vec, muladd(alpha, vec, old), muladd(alpha, vec, beta*old); split_store_block!/lanepair_store_block! take the case as a Val, same arithmetic). Tails keep axpby_at!: inside each arm LLVM folds its repeated beta test. Codegen (AVX-512, default shapes): K loops unchanged. vector_store! instructions before -> after: SIMD 1953 -> 2726, Planar 1494 -> 2004, FMAddSub 3619 -> 4757, ComplexReal/1m 3612 -> 4485, RealComplex 1900 -> 2628. Beta compares in the vector store: SIMD 96 -> 2, Planar 62 -> 13, FMAddSub/ComplexReal 167 -> 28, RealComplex 83 -> 16. Code per executed variant (beta 0/1/general): SIMD 1040/1112/1023, Planar 492/635/707, lane pairs ~1003/1291/1723, RealComplex 637/828/924. A/B vs the previous commit, medians of 5 interleaved runs: small-K (M = N = 200) Float64 K = 1: -12%, K = 2..8: -19..-27%; ComplexF64 -3..-10%; 256^3 matmul/scattered/named-kernel cases -4..+1%. Allocations 0 under --check-bounds=yes. Cost: TTFX first call Float64 3.1 -> 3.7 s, ComplexF64 2.0 -> 2.4 s (2 runs each, consistent). Full suite 31967 pass, 28m02 (previous commit 18m07; single runs on a shared workstation, but in line with the TTFX increase). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Five files instead of eight, included in this order:
- interface.jl: the contract (Microkernel, KernelMethod and pack_formats,
the add_tile/store_tile!/zero_accumulator docstrings, execute_tile!,
scale_tile!, store_prologue!, axpby helpers, shape checks).
- kernels.jl: the seven kernel types and their type-level facts (method
and layout traits, rows_per_vector, acc_index, inner, k_steps,
lanewidth, the generic constructor). The 1m and mixed-domain header
comments are folded into the kernels' docstrings.
- vecops.jl: fmaddsub, swap_pairs, addsub, deinterleave_re/_im,
interleave_planes.
- steps.jl: the K steps (scalar, real, planar, fmaddsub), the B-element
hooks, the generic add_tile and zero_accumulator.
- stores.jl: the layouts' generator-time pieces, the full-block stores and
the generic store_tile!/vector_store!/scalar_store!, plus ScalarKernel's.
Removed: types.jl, scalar.jl, simd.jl, planar.jl, onem.jl, fmaddsub.jl,
mixed.jl.
Holy traits: accumulator_layout(kernel) becomes AccumulatorLayout(kernel)
and complex_method(kernel) becomes KernelMethod(kernel), each defined on
the kernel type, with the docstring on the abstract type. The
complex_method(::Any) fallback is replaced by KernelMethod for
ScalarKernel (its only user).
Every vector kernel's inner constructor now runs the vector-shape check
(check_vector_kernel), so K{MR, NR, T, W}(descriptor) is validated and the
generic constructor only builds the descriptor. A real T for a complex
kernel is rejected by the descriptor (its message instead of the
kernel's).
Behaviour-neutral: tests 31967/31967 pass (27m59s wall); K-loop and
store instruction counts identical to cc78fd1 for all six vector
kernels; TTFX unchanged (2 interleaved runs each, before/after: load
0.5/0.5-0.6 s, first Float64 call 3.7-3.8/3.7 s, ComplexF64 2.4/2.4 s,
K=1 outer 0.6/0.6 s).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- `_sort_perm` (hand-written insertion sort) replaced by `TupleTools.sortperm`/`getindices`. TupleTools sorts by a recursive merge that takes from the right half only on a strict `lt`, so ties keep their original order, as before; allocation-free and inferred for NTuples. `import TupleTools` moves to the module file. - `_order_contract_labels` and `_choose_k_order` merged into `order_contract_labels` with an optional trailing `l2bytes` (default `nothing`: the core's L2 share, looked up only when a cost is needed). - `_l2_core_bytes` (both methods) moved to defaults.jl as `l2_core_bytes`. - conjugation.jl deleted: `_op_conjugates`/`_qs_isconj` now live in plan.jl as `op_conjugates`/`isconj(view, flag)`, both GUARDRAIL comments kept. - Names defined in labels.jl lose their leading underscore (`classify_labels`, `order_free_labels`, `leading_unit_run`, `prefer_swap`, `KOperand`, `k_operand`, `k_order_cost`, ...). The `_K_*` constants keep theirs, matching the package's other internal constants. A test helper `_macro_op_conjugates` became `_macro_conj_op`. Behaviour-neutral: K-order model, free-label order and swap rule unchanged. Per-call floor (probe_call_floor.jl, 4 interleaved runs each, medians, base 3be4a87 -> this commit, ns): F64 8x8x8 360 -> 369, 16^3 470 -> 472, d4 outer 352 -> 365; C64 8x8x8 417 -> 417, 16^3 933 -> 931, d4 outer 576 -> 577. Sample ranges overlap (F64 d4 outer 347-389 vs 347-373); 0 B/call throughout. Planning micro-benchmark with 3 K labels: K order 39 -> 18 ns, 0 B. Tests: 31967/31967 pass, 28m52s. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
`select_shape(T, method, m_length, k_length, run)` in kernel_selection.jl runs every automatic selection step in order (host shape, extent step-down, small M, `fit_to_run`), previously split between `_default_shape` in defaults.jl and `_plan_with_kernel` in plan.jl; one docstring lists the steps. defaults.jl keeps only the per-eltype cache. The mixed-method mapping now wraps the whole pipeline once. The swap decision calls it with an unbounded `k_length`, so, as before, it judges orientations without the K-dependent divisor step. `fit_to_run` merges `_store_shape` and `_demote_shape_for_run`: the MV = 4 -> MV = 2 step at any K, then the divisor step at small K, with their conditions unchanged. It now runs after the small-M step; the store step only fires on the tall shape, which the small-M step never leaves, so the order is equivalent. Also: the NEON planar override rows go (the register fit selects the same shapes; a test pins them), the dead `vb % sizeof(real(T)) == 0` conditions go, `_MixedMethod` -> `MixedMethod`, and names defined in kernel_selection.jl lose their leading underscore. The uncached `_kernel_for` and the unused `n_length` of the old pipeline are dropped. Behaviour-neutral: the automatic kernel type, shape, blocking and orientation are identical to 7032aed for Float32/64, ComplexF32/64 and both mixed pairs, m, n, k in {1,2,3,5,8,13,17,31,48,64,100,257}, four C layouts, under :avx512 (Intel and znver4), :avx2, :neon and :unknown profiles (138270 plans), and on a dense (m <= 300, every run dividing m, 7 K values) grid of the shape pipeline itself (579040 cases). Per-call floor (medians of 3 interleaved runs, ns, before -> after): Float64 8^3 372 -> 368, 16^3 484 -> 479, d4 outer 344 -> 340; ComplexF64 8^3 424 -> 415, 16^3 966 -> 963, d4 outer 580 -> 589; 0 B. TTFX: first call Float64 3.8 -> 3.8-3.9 s, ComplexF64 2.2-2.5 -> 2.4-2.5 s. Tests: 32198 pass (was 31967; the mixed-selection grid gained a K dimension), 48 min on a loaded workstation. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The unparameterised kernel type (`SIMDKernel`, `PlanarKernel`, ...) is now
the shape-free name of a complex-arithmetic scheme. `KernelMethod`, its six
singletons, `KernelMethod(kernel)` and `_kernel_type` are gone;
`default_kernel_type(T[, TA, TB])` replaces `default_method`. Menus, pack
formats, accumulator planes, blocking scaling, the shape rules, the mixed
mapping (`MixedKernel`), unpacked-B eligibility and the menu ladder dispatch
on `::Type{<:PlanarKernel}` etc.; a pair without a menu has the empty menu.
The per-scheme summary moves to the `pack_formats` docstring in kernels.jl.
ScalarKernel shares SIMDKernel's traits, so `default_blocking` works on it.
With eligibility predicted from the kernel type itself, `_hint_holds` loses
its eligibility clause.
Two places keep a type out of a value position, since a plain `Type`
regressed the per-call floor:
- the plan barrier receives `Val{<concrete kernel type>}()` (from
`menu_val`) instead of `Val(shape)` plus the method: a `Type` argument
misses the dynamic dispatch's fast cache (+350-430 ns per call measured
with the bare type);
- `select_shape` returns the kernel type as a `Val`: inference widens a type
in a returned tuple to `UnionAll`, which made `menu_val` a dynamic call
that boxed the shape (32 B per complex plan); recovering the Union by an
`isa` assertion or `===` on types cost a subtype check or structural
egal (+30-100 ns).
Behaviour-neutral: the automatic kernel type, shape, blocking and
orientation are identical to 7032aed on the same 138270-plan grid as the
previous commit, and the shape pipeline on the dense 579040-case grid.
Per-call floor (medians of 3 interleaved runs vs 7032aed, ns): Float64 8^3
372 -> 369, 16^3 484 -> 474, d4 outer 352 -> 341; ComplexF64 8^3 426 -> 422,
16^3 957 -> 954, d4 outer 586 -> 575; 0 B. TTFX: first call Float64
3.7-3.8 -> 3.7 s, ComplexF64 2.4 -> 2.3 s.
Tests: 32190 pass (redundant `KernelMethod(k) === ...` checks removed, one
ScalarKernel blocking check added), 34 min.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Blocking: `default_blocking(kernel, profile = target_profile())` evaluates the analytical model with the kernel's own tile and packed bytes per element (B sliver = NR * packed reals * sizeof(real(T))), so the "same packed byte budget" falls out of the model; undetected caches take the fallback rows scaled by packed reals as before. `_scale_blocking`, `_real_blocking_row`, the default-shape `_modelled_blocking` method and `default_blocking(::Type)` are gone. Names in blocking.jl lose their underscores (`modelled_blocking`, `fallback_blocking`, `kernel_blocking`, `roundup`). No host-defaults cache: `ResolvedDefaults`, the per-eltype slots and defaults.jl are removed. The `Val(profile.isa)` tables become functions branching on the `isa` Symbol; the small-M FMAddSub candidates are a filter over the static menu. select_shape, blocking, the K-order model, the split decision and the TensorOperations pool read `target_profile()` directly. The direct version first cost +25-57 ns per Float64 call: the `cpu_name` string compare in `profile_mv` (7 ns, three times per plan) and the divisions of `core_bytes` in the blocking model. Passing the profile down cannot help (`target_profile()` is a plain load), so instead of an eltype-keyed cache the profile itself carries the three derived values, computed by its constructor: `l2_share`, `l3_share` (the core's cache shares) and `double_pumped` (znver4/znver5). The path hint computes a kernel's default `k_block` only under a C narrower than T. Line size: `line_bytes(profile)` (detected L1d line, else 64) replaces `_K_LINE_BYTES` in the K-order model and the split decision; `_K_WALK_FAR_BYTES`/`_K_WALK_FAR_PENALTY` move to target.jl. Line packing: src/packing/line_packing.jl holds `PackSplit`, the split decision, the split enumeration and the block packer, now `pack_block_by_lines!(..., split::PackSplit)`. `pack_split(g, kgroup, spec, S, k_block, eff, rounded, k_block_requested; l2bytes = nothing)` reads the K map, tile and complexity from the operand's `SliverSpec`. src/planning/pack_split.jl is deleted. Test-only API removed: `_default_kernel(T, m, n)` and `_default_kernel(T)`; tests use `select_shape` + `kernel_from_shape` (helpers `auto_kernel`, `host_kernel`), the pool builds the host-shape kernel. Behaviour changes (equivalence grid over 7 profiles, eltypes, extents, C layouts, every menu kernel's blocking, line splits and K orders; the dense select_shape grid is identical): - Real kernels: no blocking changed on any profile. Mixed kernels: none. - Complex planar/1m/FMAddSub kernels on detected caches block by their own tile. On this host (cascadelake): planar 24x3 unchanged (96,341,786); planar 16x6 192,170,1578; 1m 12x8 120,128,2096 (was 48,341,786); FMAddSub 12x8 252,128,2096 (was 96,341,786); the AVX2 planar 4x5 160,204,1315 (was 96,341,786). Undetected caches: unchanged. - 128-byte lines (an M1-like :neon profile, and :avx512 with line 128): K orders flip on 44 grid cases, and line splits change on 22 (new splits where a line is 8 instead of 4 elements; larger L chunks). 64-byte profiles: no split or K-order decision changed. Per-call floor (tensorcontract!, 6 interleaved rounds, medians, ns, base -> new, 0 B both): Float64 8^3 362 -> 363, 16^3 473 -> 459, d4 outer 362 -> 334; ComplexF64 8^3 412 -> 405, 16^3 967 -> 948, d4 outer 570 -> 551. TTFX (3 rounds): Float64 3.8 -> 3.7 s, ComplexF64 2.4 -> 2.3 s, K=1 outer 0.6 -> 0.6 s. Tests: 32300 pass (was 32190; the per-kernel blocking test adds the rest), 54 min wall on the loaded workstation. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Same-node, interleaved per case, over every planar/1m/fmaddsub menu shape that fits the host's vectors. Validates chunk 11's change of complex blocking. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Workspace pooling is gone: the TensorOperations task-local pool, reserve!,
the reuse helpers and plan_contract's `workspace =` keyword. A plan always
builds its own workspace through the given allocator, and reusing the plan
is how buffers get reused. The two TO.tensorcontract! methods are now one:
checkpoint, then `_planned` with an `_Execute` continuation that executes
and releases the plan inside the kernel barrier (so nothing is boxed), then
reset. There is no try/finally, same as TO's own blas_contract!.
The oracle is gone too: execute_tilewise!, its tw_* buffers and the `oracle`
keyword. The tests that compared against it already compared execute!
against a dense matmul or `_lo_reference`, so those comparisons stay. The
ramp-vs-buffer test now checks against `_lo_reference`.
Workspace: the ten offset/descriptor fields become `m`/`n::GroupBuffers`
(offsets and descriptors, each a pair) and `k::NTuple{2, Vector{Int}}`. The
tile buffers are dropped; the beta-only pass uses the M/N offsets. Offsets
are plain Vectors whatever the allocator. There is one ContractWorkspace
constructor (tensoralloc for every allocator) plus `workspace_sizes`, and a
`release!(plan, allocator)` method. SlotCache/barrier_slot! move to
barrier.jl, which is now included before workspace.jl. The planning barrier
passes a fresh `Ref` because no workspace exists yet. ContractPlan documents
that it is used by one task at a time.
Tests: plan reuse with poisoned buffers is still exact, and so is the
beta-only pass. Removed the pool, reuse and reserve! tests. The TO
zero-byte-steady-state test is replaced by "the TO call allocates no more
than plan_contract" (the boxing guard). The plan_contract allocation ceiling
now adds the workspace's own bytes. Plan-reuse execute! stays 0 B.
Full suite: 32151 pass (was 32300), 20.5 min.
Cost of dropping the pool (workstation, taskset 4-7, medians; before ->
after):
call floor, DefaultAllocator, ns (bytes after; before 0 B):
Float64 8^3 336 -> 1023 (3.4 KB), 16^3 432 -> 2125 (6.9 KB),
d4 outer 323 -> 754 (2.6 KB)
ComplexF64 8^3 407 -> 1263 (4.1 KB), 16^3 940 -> 2917 (10.7 KB),
d4 outer 545 -> 1046 (2.8 KB)
@tensor n^3 matmul, DefaultAllocator, us:
Float64 32: 1.26 -> 4.2, 64: 6.7 -> 16.1, 128: 46 -> 73,
256: 382 -> 500 (0.94 MB/call)
ComplexF64 32: 5.5 -> 11.6, 64: 27 -> 45, 128: 199 -> 270,
256: 1420 -> 1600 (1.47 MB/call)
@tensor n^3 matmul, Bumper default_buffer, us:
Float64 32: 2.6 -> 2.0, 64: 10.0 -> 9.6, 128: 62 -> 61, 256: 456 -> 463
ComplexF64 32: 6.9 -> 6.7, 64: 32 -> 32, 128: 227 -> 228, 256: 1600 -> 1600
TTFX first call: Float64 3.6 -> 3.3 s, ComplexF64 2.2 -> 2.0 s.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A complex kernel keeps the k_block of real(T)'s default real kernel (B sliver at half L1d) and derives m_block/n_block from its own packed bytes at that depth, rounded down to MR/NR. The per-kernel model from chunk 11 let a complex B sliver halve k_block, which measured slower. Same-node A/B over the planar, 1m and FMAddSub menu shapes (512^3, 2048^3, 12/24x2048x2048; Genoa, Rome, Icelake): hybrid/old 0.99-1.00 (worst 1.02), per-kernel model/old up to 1.10, worst at small M. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Blocking resolution (defaults, rounding, clamping, C panel, line-by-line
packing) and path selection move before the kernel barrier, computed from the
kernel type and tile (`resolve_blocking`, `select_path`). The barrier crosses
once with the kernel (`Val{Kern}()`, or a named kernel itself), the path
singleton, the transforms and one `RefValue` holding the request plus the
resolved blocking. `ContractPlan` gains a `path` field/type parameter;
`execute!` keeps its runtime short-circuits and then calls
`execute_path!(plan, ..., plan.path)` statically, so barrier #2 is gone.
The dot path's capacity check is evaluated at plan time (`dot_fits`) from
`workspace_sizes`, which now works on the kernel type and tile.
`pack_split` takes (side, lanes, format) instead of a `SliverSpec`.
Deleted: `_path_hint`, `_default_k_block`, `_continue`, `_execute_hinted!`,
`_hint_holds`, `_splits`, `_split_path`, `_execute_across_barrier!`,
`_execute_resolved!`, `_strip_storage`, `_with_storage`, `SlotCache`,
`barrier_slot!`/`barrier_slot_slow!`, the workspace's `slots` field,
`_dot_capacity_ok`, and the global `_DOT_MODE`/`_OUTER_MODE`/
`_UNPACKED_B_MODE` Refs (now an internal `path_modes::PathModes` keyword of
`plan_contract`). `Execute` is a callable struct (execute! + release!), and
plan construction ends with `req.f(plan)`. One keyword copy method
`ContractPlan(p; ...)` replaces the spelled-out copies in the panel path and
tests. Underscores dropped on touched names; barrier.jl -> paths.jl.
Measurements (workstation, cascadelake; noisy, medians of 4 interleaved runs
of per-call medians, indicative only), before -> after:
execute! 64^3 Float64 6.8 us -> 6.8 us, 0 B
execute! 2x2 Float64 127 ns -> 85-190 ns (bimodal), 0 B
execute! 32^3 ComplexF64 5.4 us -> 5.3 us, 0 B
execute! dot 1x256x64 Float64 1.77 us -> 1.74 us, 0 B
execute! panel 32x600x32 F32/F64 30.5 us -> 30.3 us, 0 B
plan_contract 64^3 Float64 1.94 us -> 2.06 us; 72240 -> 71952 B
plan_contract 2x2 Float64 417 ns -> 405 ns; 2272 -> 1984 B
tensorcontract! 2x2 Float64 486 ns -> 484 ns; 2096 -> 1808 B
tensorcontract! 2x2 ComplexF64 687 ns -> 627 ns; 2432 -> 2144 B
TTFX (fresh process, median of 3): first Float64 plan+execute 3.44 ->
3.52 s, then ComplexF64 1.84 -> 1.93 s; first @tensor call 2.55 -> 2.53 s.
Codegen on a fixed 172-case workload (all paths, eltypes, conj, swap, split,
panel, named kernels, TensorOperations): path/nest/micro-tile/dot/outer
specializations identical (58/40/18/22/14/17/4), --trace-compile 715 -> 714
lines, except that k_length == 0 plans now compile their (never run) nest
path: +4 nests with four such cases included. Results of all 172 cases are
bitwise identical to before, with identical paths, blocking and splits.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
`NestPath{UNPACKED_B, AFF, SPLIT}` becomes `NestPath{UNPACKED_B, SPLIT}`.
Sliver axes always take the runtime `axis_of` choice (the `Val`-typed
`axis_of` methods and `throw_irregular_ramp_descriptor` go, as do the
`Val(aff_...)` arguments of `micro_tiles!` and the `AFF` lookups in
`pack_block!`). The closed-form ramp descriptors stay: `nest!` decides
per group, once per call, with `affine_ramp` on the groups it walks (a
split group is not a ramp; under `PanelPath` the plan's C maps are the
panel's dense ones; under unpacked B, N's ramp needs B's and C's maps,
as before). The 512-entry `_NEST_PATHS`/`_PANEL_PATHS` tables,
`nest_path_from_flags` and `is_ramp_map` go; `nest_path(unpacked_b,
split_a, split_b, panel)` branches over the six valid combinations, one
literal singleton per arm. `benchmark/bench_ramp_flags.jl` goes, its
question answered; `submit_ab.sh` remains for A/B runs.
Slurm evidence (same-node A/B): round 1 (7183225-7) and the unpacked-B
codegen fix (ea97e6d, 7184013-5); round 2 (7184157-9, ec3d30a vs
3b2e48e on exp/runtime-axis-types): closed-form descriptors with runtime
axis types match the static flags within noise (+-4%), except F64
partial-ramp 1.07 and ccsd(t) 1.05 on Icelake.
Equivalence vs 470f977, 743 cases: all bitwise identical; paths
identical modulo the removed parameter (base had 4 distinct AFF
patterns across its nest paths).
Specialisations on the equivalence workload, before -> after: nest!,
micro_tiles!, b_sliver 50 -> 50, execute_path! 94 -> 94, build_plan
194 -> 194, pack_block! 70 -> 72, pack_sliver!/pack! 42 -> 84,
execute_micro_tile!/execute_tile! 43 -> 160, store_tile! 32 -> 116.
The workload's plan types each met few AFF patterns, so the nest-level
counts do not drop; the leaves grow because a ramp plan now compiles
the ScatterAxis arms too (4 row/column axis-type pairs per tile).
TTFX (fresh process, medians of 3, workstation): f64 64^3 3.30 -> 5.60 s,
c64 64^3 1.88 -> 2.66 s, @tensor 1.17 -> 2.19 s, using 0.46 -> 0.55 s.
Full test suite 22m12s -> 31m50s (32160 pass).
Per-call execute! (workstation, pinned, interleaved, medians of 3 runs
of 101-sample medians; indicative), new/base: 24^3 F64 0.99, C64 1.02;
64^3 F64 0.99, C64 1.00; 12x512x512 F64 0.98, C64 0.96; 512^3 F64 1.05;
split-A (intensli_7 16^6) 0.86 (noisy); panel (F32, Float64
accumulator, 64x700x40) 1.05; ccsd(t)-like permuted F64 1.01, C64 1.01.
0 allocations in all cases.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This reverts commit b77152b. With the ramp decision at runtime every dense plan also compiles the scatter branches of the micro-tile, store and pack code (execute_tile! 43 -> 160, store_tile! 32 -> 116 specialisations): TTFX +40-90% and the suite 22 -> 32 min, for no runtime gain. The static flags pay for themselves in compile time. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
`planned` takes α and β as ordinary arguments (`nothing` = build the plan only) and carries them in `PlanRequest` across the kernel barrier instead of a continuation `f`; after the barrier `build_plan` returns the plan, or runs `execute!` and `release!` and returns nothing (`execute_request`, two methods on the field types). `Execute` goes. `plan_contract` passes `nothing`; `contract!` and `TO.tensorcontract!` convert α/β to the compute type and call `planned` directly, so `contract!` crosses one dynamic dispatch instead of two. `planned`'s options become keywords with `plan_contract`'s defaults (allocations and per-call floor unchanged). Checks: the conjugated-output check lives only in `planned`; the aliasing check moves there too, so `plan_contract`/`contract!` reject an output sharing memory with an input (new lean test). TO's arg/dim checks stay in the adapter. Renames: `_qs_labels` -> `contraction_labels`, `_qs_prepare` -> `prepare_contraction`, `_qs_check_eligible` + `_qs_throw` -> `check_strided`. The stale Zero()/One() comment now points at `static_beta`. Equivalence vs 65b3af0, 1134 cases (the existing 743: 731 plan/execute and 12 tensorcontract!; plus 3 permuted/mixed/accumulator tensorcontract! and 388 contract! calls): all bitwise identical, paths/blocking/splits identical. Per-call (workstation, pinned, interleaved base/new processes, medians of 10 runs of 101-sample medians; indicative), base -> new: tensorcontract! 2x2 F64 454 -> 457 ns, C64 698 -> 695 ns (allocs 1856/2192 B both); contract! 2x2 F64 530 -> 455 ns, 2064 -> 1856 B; contract! 64^3 F64 25.3 -> 24.5 us, 72032 -> 71824 B; execute! on a reused plan 76 -> 76 ns (2x2), 6.74 -> 6.75 us (64^3), 0 B. 64^3 calls are bimodal (~12.5 / ~19-25 us, fresh-workspace memory); one case per process, medians of 6: tensorcontract! 64^3 15.8 -> 12.6 us, contract! 64^3 21.3 -> 15.1 us, same two modes in both, 71824 B both. TTFX (fresh process, 3 interleaved runs): first plan+execute F64 64^3 3.21 -> 3.25 s, C64 1.88 -> 1.86 s, first @tensor 1.19 -> 1.13 s, first contract! (F32) 1.14 -> 1.07 s. build_plan specialisations 8 -> 8 in that session. On the equivalence workload build_plan 200 -> 304: a `contract!` call no longer shares its plan type's build_plan with `plan_contract` (as `tensorcontract!` already did not); everything below build_plan (execute_path!, nest!, packers, tiles) unchanged. Full suite: 32163 pass (19m14s); Runic clean. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
ParallelTestRunner runs every test/**/test_*.jl in its own module on a pool of worker processes; arguments select files by name prefix. Each file includes test/helpers.jl or its folder's helpers.jl, so it also runs on its own, and no file depends on another. Test files follow src/: the per-call overhead file is split into labels, axis_group, tiles, nest, execute and plan; label/K-order tests go to test_labels.jl, blocking tests to test_blocking.jl, run-length demotion to test_kernel_selection.jl, line packing to test_line_packing.jl, the C panel to test_c_panel.jl. Every @test is kept unchanged. forced_isa_runner.jl now calls Pkg.test with QS_FAKE_ISA set, and runtests.jl applies the fake target in each worker. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The audit's trims 1-9: - dot and outer paths: compared against the reference only, not also against the nest (both checks were tolerance checks). - The ScalarKernel-vs-SIMDKernel execute! testset is gone: the randomized nest test runs the same four Float64 kernels against a dense reference. - Label/K-order correctness sweeps (ccsd_t, scrambled K) run Float64 and ComplexF64; the adapter-path label test runs Float64. - The randomized nest test keeps random conj flags, blockings, alpha/beta and NaN C but no longer randomizes StridedView ops; the XOR rule stays pinned by its own testsets. - No mk_e2e runs in the planar, 1m and FMAddSub kernel files; the mixed-kernel end-to-end loop keeps ComplexReal ComplexF64 (:always, :never) and RealComplex ComplexF32, the only nest runs of their kind. - TensorOperations: @tensor/ncon and flags x op on ComplexF64 only; the StridedNative comparison uses the diagonal conj pairs for complex and (false, false) for real eltypes. - Merged duplicates: the manual-pipeline bounds/shortcut testset into mk_contract_full (destination BoundsError before any write; alpha = beta = 0 over a NaN C); the second randomized affine_ramp check; the two AxisGroup-from-labels testsets (20 trials, ranks 1-4); the 1m/FMAddSub kernel_from_shape loops into kernel_selection, their blocking ratios into test_blocking; FMAddSub's interleaved-A packing (test_pack_* cover it); the default-kernel testset into the plain-GEMM demotion check (keeping contract! == execute!); the conjugated-C nest testset (a transpose op on a complex C is now checked in test_plan); the scattered-default allocation check into test_execute, which now covers Float32 there too. Coverage, per-testset over every file (before / after): 1759 / 1759 src lines, none lost. Every eltype x kernel family pair (14) and every individual NestPath flag value per family, and per eltype and family, still reaches nest!; 11 of 86 coarse (eltype, family, flag tuple) combinations go, each recombining flag values covered elsewhere. Every shipped menu kernel compiled before is still compiled (add_tile, execute_tile!, store_tile!). Lost pack! shapes: 1e A at MR = 12 (ComplexF64) and 24 (ComplexF32), from the dropped 1m end-to-end runs, plus test-only shapes. Pkg.test: 3m38s -> 2m53s wall, 32164 -> 30888 tests; with --check-bounds=auto 2m25s -> 2m07s. Slowest files (s, before -> after): test_tensoroperations 183 -> 146, test_labels 173 -> 117, test_unpackedb 157 -> 148, test_line_packing 139 -> 129, test_onem_kernel 124 -> 115, test_dot 118 -> 60. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The 1m end-to-end runs dropped in 6b2c50d were the only tests compiling the 1e A packing at MR = 12 (ComplexF64) and 24 (ComplexF32); the fast-path test now covers every 1m menu MR instead of the menu extremes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Internal constants drop the leading underscore, like the functions in earlier chunks; _QS_ELTYPES becomes SUPPORTED_ELTYPES. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…lags Delete bench_complex_blocking.jl and bench_ramp_flags.jl (answered). The probes become the standard checks: probe_ttfx.jl times load, the first Float64/ComplexF64 plan+execute!, @tensor and contract!; probe_call_floor.jl adds contract! and reused-plan execute! (2x2, 64^3); new probe_specialisations.jl counts specialisations of the hot functions after a fixed workload. All take --label; submit_ab.sh copies a script a revision lacks from the submitting checkout. Every bench_*.jl takes --smoke (helpers in harness.jl); blocking-model-only helpers move into that script. Fix the blocking model summary passing a dtype to named_rows. Add benchmark/README.md. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
plot_bench_to_suite.jl now draws two figures with one panel per dtype: throughput against arithmetic intensity (colour = total work, filled = QuasiStrided, open = StridedBLAS, the two joined per case) and the QuasiStrided/StridedBLAS time ratio against total work (colour = intensity, marker = category), with the geomean and faster count in the panel title and the three most extreme cases at each end numbered and named below. --dtypes, --categories and --tags filter the rows; the filters name the output files. The per-group figures and the synthetic violins are gone. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
plot_bench_to_suite.jl draws a single figure: one column per dtype, three rows on a shared arithmetic-intensity axis (QuasiStrided GFLOP/s, StridedBLAS GFLOP/s on the same scale, and the time ratio with the geomean and faster count). Colour = total work on a trimmed batlow map with one colorbar, marker = category, semi-transparent markers; the pairing segments and the extreme-case labels are gone. Log axes get plain tick labels spaced to their span. Output: bench_to_suite_overview[_<filters>].png. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
One row per dtype, columns QuasiStrided GFLOP/s | StridedBLAS GFLOP/s | time ratio on the linked arithmetic-intensity axis; the two throughput panels of a row share y-limits. The geomean and faster count move into a boxed label in the ratio panel's corner. Without --dtypes the plot shows whichever of Float64 and ComplexF64 the CSV has, else every dtype in it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The geomean box sat over the lowest ratios and could hide an outlier; the y-range now leaves an empty band below the data for it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Member
Author
Upstream suite after the cleanup
The network cases, which an earlier complex-only run (older code) showed at about 1.1–1.2×, are now clearly faster. The remaining losses are per-call overhead on batched and tiny BLAS-shaped contractions: #23. 🤖 Generated with Claude Code |
lkdvos
added a commit
that referenced
this pull request
Oct 7, 2026
At 16 ymm registers LLVM's haswell/znver2 scheduler hoists every B broadcast of a planar or fmaddsub K step to the top of the loop body, holds them all live and spills them. `kstep_fence()`, an empty `asm sideeffect` with a memory clobber (no instruction), after each column group pins the loads next to their FMAs; the FMAs stay free to move. On AVX2, fmaddsub also runs OpenBLAS's Zen zgemm order: the swap(a)*bi pass, a fence, then the a*br pass on A loaded again, so a and swap(a) are never live together. Only 256-bit kernels (`W * sizeof(R) == 32`) take the fenced body; every other shape runs the unchanged generator. code_llvm/code_native of `add_tile` are identical to ad94b73 for all 57 non-256-bit complex and mixed-domain menu shapes (B packed and B in place) under -C icelake-server, znver4, znver2 and haswell. Hot-loop stack references per K step, Julia 1.12.7, -C znver2 (haswell identical), ad94b73 -> this commit, B packed / B read in place (the #22 unpacked-B codegen does not change the picture): planar 4x5/W4, 8x5/W8 4 / 4 -> 0 / 0 fmaddsub 4x6/W4, 8x6/W8 24 / 24 -> 0 / 0 fmaddsub 4x5/W4, 8x5/W8 0 / 0 -> 0 / 0 planar 4x6/W4, 8x6/W8 14 / 12 -> 0 / 13 (not in any menu) The mixed-domain kernels run the real SIMDKernel step and do not spill. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
lkdvos
added a commit
that referenced
this pull request
Oct 7, 2026
… vector store On AVX2, `select_shape` hands a complex contraction to FMAddSub 4x6/W4 (ComplexF64) or 8x6/W8 (ComplexF32) when C's unit run along M fills whole `MR`-row slivers (`run == m_length || run % MR == 0`); otherwise it keeps planar 4x5/8x5. FMAddSub's single accumulator plane holds 12 accumulators + 2 A + 1 broadcast, 15 of 16 registers, and stays spill-free with B packed or in place. Planar NR = 6 fills all 16 and spills 13 per K step with B in place, so it is not a candidate. Where C's rows defeat the vector store, fmaddsub's lane-pair scalar store costs about 1.3x planar's per element, and with a short K the store dominates (ccsd_t_2/3/4, ao2mo_2/3), hence the run condition. `plan_contract` now computes the run for complex `T` too. Plans over the 490 suite :contract cases with uniform legs, same-eltype and both mixed-domain orientations, icelake-server/znver4/znver2 profiles, ad94b73 vs this branch: identical everywhere except 419 same-eltype plans on znver2, which change kernel (planar -> fmaddsub) and the blocking derived from it; the execution path is unchanged. Rome, same-node A/B of the same change on 4a77066, before the cleanup (#22) (jobs 7143376/7143377, two rounds each; GF/s of the default kernel, main planar 4x5 -> fmaddsub 4x6, OpenBLAS in brackets): ComplexF64 64^3 35.9 -> 40.6 [42.3], 256^3 42.1 -> 44.6 [49.3], 512^3 41.2 -> 48.2 [50.3], 2048^3 40.5 -> 45.1 [51.2], smallMN 29.3 -> 37.0 [30.2]; time geomean 0.87, no case slower ComplexF32 64^3 61.1 -> 76.7 [72.5], 256^3 80.7 -> 90.4 [95.2], 2048^3 81.5 -> 99.3 [102.8], smallN 44.0 -> 73.0 [53.1]; time geomean 0.79, no case slower Suite (complex :contract and :network): QuasiStrided time geomean 0.78 (ComplexF32) / 0.83 (ComplexF64) of main. Consistently slower: the dim32/96_1_3_1 contract_scrambled cases (1.23-1.47), and in ComplexF64 intensli_7_dim24 and dim8-16 0_2_3 cases (1.06-1.14). A diagnostic run traced the scrambled cases to the kernel method (fmaddsub), not the fence. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
lkdvos
added a commit
that referenced
this pull request
Oct 8, 2026
At 16 ymm registers LLVM's haswell/znver2 scheduler hoists every B broadcast of a planar or fmaddsub K step to the top of the loop body, holds them all live and spills them. `memory_fence()`, an empty `asm sideeffect` with a memory clobber (no instruction), after each column group pins the loads next to their FMAs; the FMAs stay free to move. On AVX2, fmaddsub also runs OpenBLAS's Zen zgemm order: the swap(a)*bi pass, a fence, then the a*br pass on A loaded again, so a and swap(a) are never live together. Only 256-bit kernels (`W * sizeof(R) == 32`) take the fenced order; every other shape gets the same statements in the same order as before. code_llvm/code_native of `add_tile` are identical to ad94b73 for all 57 non-256-bit complex and mixed-domain menu shapes (B packed and B in place) under -C icelake-server, znver4, znver2 and haswell. Hot-loop stack references per K step, Julia 1.12.7, -C znver2 (haswell identical), ad94b73 -> this commit, B packed / B read in place (the #22 unpacked-B codegen does not change the picture): planar 4x5/W4, 8x5/W8 4 / 4 -> 0 / 0 fmaddsub 4x6/W4, 8x6/W8 24 / 24 -> 0 / 0 fmaddsub 4x5/W4, 8x5/W8 0 / 0 -> 0 / 0 planar 4x6/W4, 8x6/W8 14 / 12 -> 0 / 13 (not in any menu) The mixed-domain kernels run the real SIMDKernel step and do not spill. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
lkdvos
added a commit
that referenced
this pull request
Oct 8, 2026
… vector store On AVX2, `select_shape` hands a complex contraction to FMAddSub 4x6/W4 (ComplexF64) or 8x6/W8 (ComplexF32) when C's unit run along M fills whole `MR`-row slivers (`run == m_length || run % MR == 0`); otherwise it keeps planar 4x5/8x5. FMAddSub's single accumulator plane holds 12 accumulators + 2 A + 1 broadcast, 15 of 16 registers, and stays spill-free with B packed or in place. Planar NR = 6 fills all 16 and spills 13 per K step with B in place, so it is not a candidate. Where C's rows defeat the vector store, fmaddsub's lane-pair scalar store costs about 1.3x planar's per element, and with a short K the store dominates (ccsd_t_2/3/4, ao2mo_2/3), hence the run condition. `plan_contract` now computes the run for complex `T` too. The FMAddSub NR = 5 entries (4,5,4)/(8,5,8) leave the menus: nothing selects them, and the fenced NR = 6 tiles are spill-free. One `shape_override` row now gives both lane-pair kernels (1m, FMAddSub) their AVX2 shape; for OneMKernel ComplexF32 that is (8,6,8) where the fit used to hand it the AVX-512 menu head (24,8,16). Plans over the 490 suite :contract cases with uniform legs, same-eltype and both mixed-domain orientations, icelake-server/znver4/znver2 profiles, ad94b73 vs this branch: identical everywhere except 419 same-eltype plans on znver2, which change kernel (planar -> fmaddsub) and the blocking derived from it; the execution path is unchanged. Rome, same-node A/B of the same change on 4a77066, before the cleanup (#22) (jobs 7143376/7143377, two rounds each; GF/s of the default kernel, main planar 4x5 -> fmaddsub 4x6, OpenBLAS in brackets): ComplexF64 64^3 35.9 -> 40.6 [42.3], 256^3 42.1 -> 44.6 [49.3], 512^3 41.2 -> 48.2 [50.3], 2048^3 40.5 -> 45.1 [51.2], smallMN 29.3 -> 37.0 [30.2]; time geomean 0.87, no case slower ComplexF32 64^3 61.1 -> 76.7 [72.5], 256^3 80.7 -> 90.4 [95.2], 2048^3 81.5 -> 99.3 [102.8], smallN 44.0 -> 73.0 [53.1]; time geomean 0.79, no case slower Suite (complex :contract and :network): QuasiStrided time geomean 0.78 (ComplexF32) / 0.83 (ComplexF64) of main. Consistently slower: the dim32/96_1_3_1 contract_scrambled cases (1.23-1.47), and in ComplexF64 intensli_7_dim24 and dim8-16 0_2_3 cases (1.06-1.14). A diagnostic run traced the scrambled cases to the kernel method (fmaddsub), not the fence. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
lkdvos
added a commit
that referenced
this pull request
Oct 8, 2026
…k/store (#16) * Fence the AVX2 complex K steps so their broadcasts stop spilling At 16 ymm registers LLVM's haswell/znver2 scheduler hoists every B broadcast of a planar or fmaddsub K step to the top of the loop body, holds them all live and spills them. `memory_fence()`, an empty `asm sideeffect` with a memory clobber (no instruction), after each column group pins the loads next to their FMAs; the FMAs stay free to move. On AVX2, fmaddsub also runs OpenBLAS's Zen zgemm order: the swap(a)*bi pass, a fence, then the a*br pass on A loaded again, so a and swap(a) are never live together. Only 256-bit kernels (`W * sizeof(R) == 32`) take the fenced order; every other shape gets the same statements in the same order as before. code_llvm/code_native of `add_tile` are identical to ad94b73 for all 57 non-256-bit complex and mixed-domain menu shapes (B packed and B in place) under -C icelake-server, znver4, znver2 and haswell. Hot-loop stack references per K step, Julia 1.12.7, -C znver2 (haswell identical), ad94b73 -> this commit, B packed / B read in place (the #22 unpacked-B codegen does not change the picture): planar 4x5/W4, 8x5/W8 4 / 4 -> 0 / 0 fmaddsub 4x6/W4, 8x6/W8 24 / 24 -> 0 / 0 fmaddsub 4x5/W4, 8x5/W8 0 / 0 -> 0 / 0 planar 4x6/W4, 8x6/W8 14 / 12 -> 0 / 13 (not in any menu) The mixed-domain kernels run the real SIMDKernel step and do not spill. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Open the complex pack and store fast paths to AVX2 `complex_fastpath_isa_eligible` was AVX-512 only, so on AVX2 every complex sliver packed through the scalar loop and every complex tile (planar, fmaddsub and the mixed-domain kernels, which share those stores) stored element by element. The AVX2 codegen is the AVX-512 path at half width. NEON and unknown targets keep the scalar loops. Measured on the pre-refactor branch (Cascade Lake, QS_FAKE_ISA=avx2 -C znver2, gate off -> on, GF/s): kernel-only planar 8x6/W8 85.4 -> 108.5; engine ComplexF32 2000^3 94.2 -> 105.1, shallowK 47.0 -> 85.5, smallN 47.8 -> 70.4. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * AVX2 complex default: fmaddsub 4x6/W4, 8x6/W8 where C's rows take its vector store On AVX2, `select_shape` hands a complex contraction to FMAddSub 4x6/W4 (ComplexF64) or 8x6/W8 (ComplexF32) when C's unit run along M fills whole `MR`-row slivers (`run == m_length || run % MR == 0`); otherwise it keeps planar 4x5/8x5. FMAddSub's single accumulator plane holds 12 accumulators + 2 A + 1 broadcast, 15 of 16 registers, and stays spill-free with B packed or in place. Planar NR = 6 fills all 16 and spills 13 per K step with B in place, so it is not a candidate. Where C's rows defeat the vector store, fmaddsub's lane-pair scalar store costs about 1.3x planar's per element, and with a short K the store dominates (ccsd_t_2/3/4, ao2mo_2/3), hence the run condition. `plan_contract` now computes the run for complex `T` too. The FMAddSub NR = 5 entries (4,5,4)/(8,5,8) leave the menus: nothing selects them, and the fenced NR = 6 tiles are spill-free. One `shape_override` row now gives both lane-pair kernels (1m, FMAddSub) their AVX2 shape; for OneMKernel ComplexF32 that is (8,6,8) where the fit used to hand it the AVX-512 menu head (24,8,16). Plans over the 490 suite :contract cases with uniform legs, same-eltype and both mixed-domain orientations, icelake-server/znver4/znver2 profiles, ad94b73 vs this branch: identical everywhere except 419 same-eltype plans on znver2, which change kernel (planar -> fmaddsub) and the blocking derived from it; the execution path is unchanged. Rome, same-node A/B of the same change on 4a77066, before the cleanup (#22) (jobs 7143376/7143377, two rounds each; GF/s of the default kernel, main planar 4x5 -> fmaddsub 4x6, OpenBLAS in brackets): ComplexF64 64^3 35.9 -> 40.6 [42.3], 256^3 42.1 -> 44.6 [49.3], 512^3 41.2 -> 48.2 [50.3], 2048^3 40.5 -> 45.1 [51.2], smallMN 29.3 -> 37.0 [30.2]; time geomean 0.87, no case slower ComplexF32 64^3 61.1 -> 76.7 [72.5], 256^3 80.7 -> 90.4 [95.2], 2048^3 81.5 -> 99.3 [102.8], smallN 44.0 -> 73.0 [53.1]; time geomean 0.79, no case slower Suite (complex :contract and :network): QuasiStrided time geomean 0.78 (ComplexF32) / 0.83 (ComplexF64) of main. Consistently slower: the dim32/96_1_3_1 contract_scrambled cases (1.23-1.47), and in ComplexF64 intensli_7_dim24 and dim8-16 0_2_3 cases (1.06-1.14). A diagnostic run traced the scrambled cases to the kernel method (fmaddsub), not the fence. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
A chunk-by-chunk review of the whole package, aimed at a codebase that is easier to read and maintain. Each chunk was explained, discussed and then changed in its own commit(s), verified with the full suite, bitwise old-vs-new equivalence scripts where behaviour had to stay the same, per-call floor and allocation checks, TTFX and specialisation counts. Performance claims come from same-node Slurm A/B runs.
Breaking changes (0.1, unregistered)
plan_contract'smc/kc/nckeywords and theBlockingfields are nowm_block/k_block/n_block;mr(k)/nr(k)aretile_size(k)/tile_size(k, i)(now public, withsliver_width).plan_contractandcontract!now also reject an output that might alias an input, as the TensorOperations backend already did (coarse for now: any shared parent; see Precise aliasing check between output and inputs #18).What changed, by area
target.jl). Own detection kept;TargetProfilederives vector width, register count, per-core cache shares and the cache-line size at construction.AxisGroupconstructor from labels; one-based coordinates inside tiles, zero-based addresses; oneTiletype withgetindex/setindex!;ScatterAxisis pointer-backed so the axis union stays unboxed.pack!(panel, tile, sliver_spec(kernel, i), transform)with A and B on equal footing (B packed astranspose(tile)); line packing inpacking/line_packing.jl;UnpackedBViewis a B source next toPackedPanel.ComplexMethodtraits are gone: the unparameterised kernel type (e.g.PlanarKernel) names a scheme. A kernel is a K step plus an accumulator layout (real, split, lane pairs); one genericadd_tile/store_tile!; β cases dispatch on VectorInterface'sZero/Oneat the existing branch points, with the β hoist kept. Unpacked-B codegen fix: dense B loads through per-column pointers and no unrolling of that K loop (ComplexF64 small M 2× faster on Icelake, Float64 small M/64³ 16–27% faster on Genoa, Rome neutral).select_shapepipeline; analytical blocking with the hybrid complex rule (k_blockfrom the real default kernel's B sliver,m_block/n_blockfrom the kernel's own bytes); no workspace pool and no tile-wise oracle (reuse a plan to reuse buffers).execute!on a reused plan has no dynamic dispatch. The hint machinery, storage stripping and the workspace's slot cache are gone. The static ramp flags stay: a runtime-ramp variant was measured (same runtime, but TTFX +40–90% and the suite 22 → 32 min) and reverted.execute.jlsplit intoexecute.jl,nest.jlandc_panel.jl; trivialEmptyPath/ScalePathfor empty extents and K = 0; one tile loop for packed and unpacked B; the dot path dispatches on its orientation and the workspace is sized per path (dot plans allocate 8.5 KB instead of 150 KB).contract!goes through one dynamic crossing instead of two (2×2: 530 → 455 ns).src/, per-folder helpers, Aqua ambiguity checks for the package itself. A coverage audit drove trims that lose nosrcline and keep every eltype × kernel family and path flag reaching the nest.Pkg.test22 min → under 3 min (--check-bounds=yeskept;julia_args = ["--check-bounds=auto"]is the fast local run).submit_ab.sh);--smokeeverywhere;benchmark/README.md; a new suite overview plot (one row per eltype: QuasiStrided GF/s, StridedBLAS GF/s, time ratio, against arithmetic intensity, coloured by total work, with tag filters).Follow-ups
🤖 Generated with Claude Code