Skip to content

amd64: preserve XMM/YMM/ZMM register aliasing and upper-bit semantics - #45

Merged
xushiwei merged 4 commits into
xgo-dev:mainfrom
zhouguangyuan0718:codex/amd64-vector-register-alias-20261004
Oct 5, 2026
Merged

xushiwei merged 4 commits into
xgo-dev:mainfrom
zhouguangyuan0718:codex/amd64-vector-register-alias-20261004

Conversation

@zhouguangyuan0718

@zhouguangyuan0718 zhouguangyuan0718 commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor

Go's amd64 CRC32 AVX512 path mixes XMM and ZMM views of the same register. The translator currently allocates independent XMM/YMM/ZMM storage, so writes such as VMOVAPS X0, X10 are invisible to a later read of Z10, and a ZMM result is invisible through its XMM view. This produces incorrect IEEE CRC32 values for buffers of at least 1024 bytes and causes gzip checksum failures when the path is enabled, as observed in xgo-dev/llgo#2722.

Allocate one backing slot per register number at the widest width used by the function, and share its low bytes across XMM/YMM/ZMM views. Preserve upper bytes for legacy SSE writes, clear them for VEX/EVEX writes, and implement VZEROUPPER/VZEROALL for registers 0–15 instead of treating them as no-ops.

Refresh the strict latest-version corpus pins to compress v1.20.1 and lz4 v4.1.33, including the new lz4 amd64 assembly inventory. The updated lz4 ARM64 decoder uses FSTPQ; implement full-width register-pair stores with offset, pre-indexed, post-indexed, and global-symbol addressing.

Validation:

  • Six executable regressions cover wide-to-narrow visibility, legacy SSE upper-byte preservation, VEX 128/256-bit upper-byte clearing, and both zeroing instructions. All six fail on the previous implementation and pass with this change.
  • go test ./... -count=1 passed on macOS arm64 with LLVM 22.1.8 and Go 1.27.0; the amd64 regression executables ran through Rosetta.
  • An LLGo CRC probe explicitly selected the translated AVX512 branch and compared it with an independent scalar polynomial reference. The previous translator returned incorrect values at 1024, 2048, and 4096 bytes; all sizes passed after this fix. This validates the translated branch through baseline LLVM legalization, not native AVX512 hardware execution.
  • LLGo CRC32 and binary-fixture suites passed with the patched dependency, including large-buffer, alignment, and initial-CRC coverage.
  • The new ARM64 pair-store fixture passes the native Go assembler oracle. LLVM translation fails before the FSTPQ implementation and passes with it; the LLVM executable verifies both 128-bit values, surrounding memory, and base-register writeback.
  • Both updated corpus suites pass strict latest-version and inventory checks across 11 targets: compress 42/42 and lz4 17/17 assembly translations and LLVM object compilations.
  • scripts/check-arm64-plan9-corpus.sh passed after inspecting the before/after form reports and updating the baseline for exactly two newly supported FSTPQ forms. FSTPQ is also included in the native-Go-accepted family check (501 accepted instructions, zero unsupported forms).
  • Official Go assembler coverage checks pass locally with all eight Go versions (1.20–1.27), covering 386, amd64, ARM, ARM64, and Wasm. Before/after report comparisons confirm exactly five newly supported ARM64 FSTPQ forms and two additional executable-conformance forms; other architectures and encoder coverage are unchanged.

@codecov

codecov Bot commented Oct 4, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 92.18750% with 5 lines in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
arm64_lower_vec.go 78.26% 5 Missing ⚠️

📢 Thoughts on this report? Let us know!

@cpunion

cpunion commented Oct 4, 2026

Copy link
Copy Markdown
Collaborator

Reviewed commit 4a3b0bb07afc0976de76b5b3824be069cb3015f5. No blocking issues found in this change.

The shared backing slots model XMM/YMM/ZMM aliasing correctly. Legacy SSE writes preserve the upper bytes, VEX/EVEX writes clear them, and VZEROUPPER/VZEROALL leave registers 16–31 unchanged.

Validation:

  • The root module's go test ./... -count=1 passed, including the ARM64 pair-store conformance tests.
  • The six new aliasing regressions passed.
  • Twelve additional local boundary probes passed, covering XMM/YMM views, identical source/destination registers, memory stores, EVEX narrow writes, and preservation of Z16/Z31 across VZERO*.

Non-blocking suggestion: add the Z16/Z31 preservation cases to the permanent regression suite. This boundary is explicitly specified in the Intel SDM.

The x86 execution checks used Rosetta and LLVM legalization; they do not establish native AVX512 hardware coverage.

@xushiwei
xushiwei merged commit e2164ac into xgo-dev:main Oct 5, 2026
56 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants