Skip to content

feat(cpu): ARM64 NEON GEMM micro-kernel and element-wise ops (#106) - #152

Merged
kolkov merged 1 commit into
mainfrom
feat/arm64-neon-sprint4
Aug 4, 2026
Merged

kolkov merged 1 commit into
mainfrom
feat/arm64-neon-sprint4

Conversation

@kolkov

@kolkov kolkov commented Aug 4, 2026

Copy link
Copy Markdown
Contributor

Summary

ARM64 NEON SIMD support — GEMM micro-kernel and element-wise float32 operations.

Phase 1: GEMM micro-kernel (Plan 9 mnemonics)

  • 4×8 register-blocked GEMM (gemmMicroKernel4x8NEON): 8 accumulator V-registers, VFMLA inner loop, packed A/B panels. Mirrors the amd64 AVX2 6×16 kernel pattern.
  • 1×8 remainder kernel (gemmMicroKernel1x8NEON): for row tail and GEMV shapes.
  • Runtime dispatch via cpu.ARM64.HasASIMD (mandatory on all arm64 hardware).
  • Full Go dispatch layer: packing (packA4, packB8), tile/tail/GEMV paths.

Phase 2: Element-wise ops (WORD encoding)

  • 8 kernels: Add/Sub/Mul/Div × (inplace + vectorized)
  • WORD-encoded FADD/FSUB/FMUL/FDIV for .4S vectors (not yet Plan 9 mnemonics in Go 1.26)
  • Scalar tail for remainder elements
  • Wired into existing simdXxxFloat32 function pointers with simdMinLen=32 threshold

Files

  • gemm_microkernel_arm64.s — 152 lines Plan 9 assembly
  • gemm_microkernel_stub_arm64.go — Go declarations
  • matmul_gemm_arm64.go — dispatch + packing (197 lines)
  • matmul_gemm_arm64_test.go — 34 shapes + dispatch + allocs + benchmark
  • simd_neon_arm64.s — 352 lines (8 kernels)
  • simd_neon_stub_arm64.go — Go declarations
  • simd_neon_arm64.go — init() registration
  • matmul_gemm.go — gemmMinCols → var (8 for arm64, 16 for amd64)

Test plan

  • go build ./... (amd64)
  • GOARCH=arm64 GOOS=linux go build ./...
  • GOARCH=arm64 GOOS=darwin go build ./...
  • GOOS=js GOARCH=wasm go build ./...
  • go vet ./...
  • gofmt — clean
  • golangci-lint run — 0 issues
  • go test -short ./... — 22/22 pass, 0 failures
  • ARM64 tests run on CI (macos-latest = Apple Silicon)

Closes #106

Phase 1: 4×8 register-blocked GEMM micro-kernel using VFMLA/VLD1/VLD1R/VST1
Plan 9 mnemonics. 8 accumulator V-registers, packed A/B panels matching the
amd64 AVX2 GEMM pattern. gemmMr=4, gemmNr=8. Runtime dispatch via HasASIMD.

Phase 2: Element-wise Add/Sub/Mul/Div float32 using WORD-encoded FADD/FSUB/
FMUL/FDIV (.4S vectors). Scalar tail for remainder. All 8 kernels wired into
existing simd function pointers. simdMinLen=32 threshold applies.

Cross-compile verified: arm64/linux, arm64/darwin, amd64, wasm.
34 GEMM test shapes + dispatch + allocs + benchmark. Lint: 0 issues.
@codecov

codecov Bot commented Aug 4, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@kolkov
kolkov merged commit 6e72d1d into main Aug 4, 2026
11 checks passed
@kolkov
kolkov deleted the feat/arm64-neon-sprint4 branch August 4, 2026 13:43
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

feat: ARM64 NEON SIMD kernels (GEMM, element-wise ops)

1 participant