Skip to content

simd: expand SIMD128 operations and runtime integration - #2722

Merged
xushiwei merged 20 commits into
xgo-dev:mainfrom
zhouguangyuan0718:codex/simd-next-20261002
Oct 6, 2026
Merged

xushiwei merged 20 commits into
xgo-dev:mainfrom
zhouguangyuan0718:codex/simd-next-20261002

Conversation

@zhouguangyuan0718

@zhouguangyuan0718 zhouguangyuan0718 commented Oct 3, 2026 •

Copy link
Copy Markdown
Collaborator

SIMD128 programs can now load and store arrays/slices, perform arithmetic and conversions, compare/select lanes, and use integer shifts, saturation, lookups, and permutations across the applicable amd64, arm64, and Wasm APIs.

The lowering keeps LLVM vector values across calls and computations while preserving Go aggregate storage and layout. It handles bounds/nil checks, shift limits, NaNs, signed zero, conversion overflow, and architecture-specific floating min/max behavior. amd64 permutations stay within the compilation baseline: constant indices fold to vector shuffles at O2, while dynamic indices retain a per-lane fallback.

Runtime and build integration:

  • Initialize effective CPU features before archsimd and user initialization, including the official GODEBUG=cpu.* policy and platform hooks.
  • Attach source patches to every test-package variant.
  • Bridge vector arguments/results through memory at JavaScript SjLj boundaries for both GoJS and Emscripten. This fixes the default GoJS O0 recovery boundary previously failing with TypeError: type incompatibility when transforming from/to JS.
  • Preserve ordinary vector calls and prevent later backend/LTO inlining from moving unbridged calls into recovery functions. Omit bridges and these inlining restrictions when both compilation and linking select native Wasm SjLj. Determine that mode from effective driver arguments and EMCC_CFLAGS; Memory64's explicit JS SjLj codegen retains the bridge.
  • Add executable tests for memory, arithmetic, masks, conversions, cross-package/interface/closure calls, defer/recover, goroutines, and a hex-encoding correctness workload.
  • Restrict LookupOrZero lowering to arm64/wasm; unsupported targets use the existing intrinsic fallback.
  • Use released plan9asm v0.6.2 (amd64: preserve XMM/YMM/ZMM register aliasing and upper-bit semantics plan9asm#45) for shared XMM/YMM/ZMM register storage. Add scalar-reference CRC folding coverage and Windows CPU override expectations matching its environment-initialization order.

Rebased onto ec9c2488b, including the current Wasmer runner, WASI LTO, and Emscripten native EH capability support. The SIMD lowering and CPU/build integration remain necessary on that baseline. The README now includes WASI Thin/Full LTO commands and GoJS boundary coverage.

Validation of the baseline refresh and SjLj fix (Go 1.27.0, LLVM 22.1.8, Emscripten 6.0.8, Node 24.19.0, Wasmer 7.5.0):

  • Full internal/clang and internal/crosscompile tests passed.
  • Focused SIMD, CPU initialization, and Emscripten tests in cl, ssa, and internal/build passed. Bridge regressions check GoJS and Emscripten, unchanged IR under native SjLj, late inlining, flag precedence, C++ link-driver arguments, and Memory64's retained JS codegen.
  • O0 boundary executables passed for both GoJS and Emscripten with default JavaScript SjLj and explicit native Wasm SjLj. Execution checks bypassed the build cache.
  • GoJS boundary executables passed at O2 with requested Thin/Full LTO options in both JavaScript and native Wasm SjLj modes, bypassing the build cache. Follow-up driver inspection in wasm: unify browser exception handling on native SjLj #2734 found that browser targets did not forward those LTO options to the compiler/linker, so these earlier runs do not qualify actual browser LTO. wasm: unify browser exception handling on native SjLj #2734 fixes that forwarding and validates actual LTO separately; its integration checkout with this PR also passes the full SIMD suite with the entire JS bridge pass disabled.
  • Full WASI SIMD tests passed at O2 Thin and Full LTO with Wasmer 7.5.0.

Earlier validation retained from this PR:

  • Full cl, ssa, internal/packages, and internal/plan9asm tests; native arm64 SIMD at O0/O2 and Full LTO; amd64 SIMD via Rosetta at GOAMD64=v1 with CPU overrides; WASI SIMD at O0/O2; full Emscripten SIMD at O2; official Go native and WASI shared behavior tests.
  • Forced translated CRC folding matched the scalar reference at 64, 512, 1024, 2048, and 4096 bytes, with the wide path enabled and disabled. This used baseline LLVM legalization through Rosetta, not native AVX512 hardware.
  • Native CRC, binary-parser fixture, and CPU profile SIGPROF tests; amd64 CRC and binary-parser fixture tests via Rosetta.

Limits: dynamic amd64 permutation performance remains future work; no hardware-specific dispatch is added. The full O0 Emscripten test executable exceeds Node's local-variable limit, so O0 uses the smaller boundary executable. Official Go 1.27.0 WASI intrinsics do not implement LLGo's nil-array panic behavior; nil assertions remain enabled for every LLGo target and official native Go. Full standard-library internal/cpu tests still encounter the separately reproduced duplicate-package-symbol issue. SIMD reflection, portable simd specialization, general FMV, and 256/512-bit vectors remain outside this stage.

Follows #2581; implements additional operations from #2568.

@codecov

codecov Bot commented Oct 3, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 96.18834% with 17 lines in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
cl/simd.go 89.15% 9 Missing ⚠️
internal/build/wasm_simd_calls.go 97.45% 3 Missing ⚠️
ssa/simd.go 97.65% 3 Missing ⚠️
internal/clang/clang.go 0.00% 1 Missing ⚠️
internal/plan9asm/translate.go 80.00% 1 Missing ⚠️

📢 Thoughts on this report? Let us know!

@fennoai fennoai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

SIMD128 implementation review

Reviewed the SIMD128 lowering, Emscripten SjLj bridging, CPU-detection patches, and GOAMD64 wiring across the 783ec4f...a09e36e range. The implementation is strong overall: poison-avoidance (shift clamping, permute index masking, saturating fptosi.sat/fptoui.sat, compile-time lane-index checks), nil-deref guards on load/store, and the "vector-bits-only" GC reasoning for the Emscripten stack slots all check out. The package-dedup fix (keying checked by pkg.ID instead of PkgPath) and the single non-cumulative GOAMD64_vN macro injection both match the intended semantics.

A few mostly-minor items are noted inline. Nothing blocking.

Notes considered but not flagged:

  • The empty-slice/NUL-termination assumptions in sysctl_darwin_llgo.go are carried over verbatim from the deleted os_darwin.go, fed only by fixed stdlib string literals — preserved behavior, not a regression.
  • requireSIMDFeatures/finishSIMDFeatures cost is pre-existing (unchanged by this PR), so not in scope here.

Comment thread test/simd/README.md Outdated

The current WAMR 2.4.5 classic-interpreter profile rejects `v128` function
types with `unknown value type`, even though its build reports SIMD enabled.
The WASI suite is therefore blocked at module loading; the same failure is

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] README says WASI suite is blocked, but CI still runs it

This states the WASI suite is "blocked at module loading" on the WAMR classic-interpreter profile. However the Running section just above (llgo test -O2 -target wasi -emulator ... ./test/simd/...) and CI (.github/workflows/llgo.yml:898-901) actively run exactly that LLGo WASI SIMD invocation. If the v128 module-load failure is real on the shipped WAMR profile, that command/CI step would fail at load. Please reconcile: either the suite does run (clarify under which engine) or the command/CI step should be marked skipped/expected-fail.

Comment thread ssa/simd_permute.go
n := x.ll.VectorSize()
mask := llvm.ConstInt(indices.ll.ElementType(), uint64(n-1), false)
result := llvm.Undef(x.ll)
for i := 0; i < n; i++ {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Dynamic Permute emits a scalarized per-lane gather

For SIMDPermute/SIMDPermuteOrZero the fallback emits, per lane, extractelement(index) + and + extractelement(value) + optional icmp/select + insertelement — a serial ~4-6 instruction chain through result (≈80-100 IR instructions for 16 lanes). LookupOrZero already uses the hardware wasm.swizzle/neon.tbl1 intrinsic, but the dynamic-index Permute path has none, so unless the backend recollapses this it is a runtime SIMD-quality regression. If indices are commonly compile-time constants, a single shufflevector would be far better. Worth confirming the optimizer folds this back to a vector shuffle.

Comment thread ssa/simd.go Outdated
SIMDTrunc: "llvm.trunc", SIMDRound: "llvm.roundeven",
}

func (b Builder) simdFeatures(op SIMDOp) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P3] simdFeatures ignores its op parameter

func (b Builder) simdFeatures(op SIMDOp) never uses op; it only forwards to requireSIMDFeatures(). This implies per-op feature logic that no longer exists. Consider dropping the parameter (or the wrapper) and calling requireSIMDFeatures() directly.

Comment thread ssa/simd.go Outdated
info := simdLanes(x.RawType()).Elem().Underlying().(*types.Basic).Info()
var cond llvm.Value
if info&types.IsFloat != 0 {
pred := map[SIMDOp]llvm.FloatPredicate{

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P3] Per-call predicate-map allocation in simdCompare

simdCompare allocates and populates a fresh map[SIMDOp]llvm.FloatPredicate (or Int) literal on every comparison lowered, just for one lookup; simdIntegerIntrinsic (simd_integer.go:60) does the same with a map[SIMDOp]string. A package-level var (as already done for simdFloatUnary) or a switch avoids the per-call allocation. Minor.

Comment thread ssa/simd_convert.go Outdated
v = b.impl.CreateBitCast(x.impl, typ, "")
}
if n != result.ll.VectorSize() {
// Narrowing conversions leave the unused upper lanes zero on these targets.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P3] "Narrowing conversions" comment conflates element vs lane count

This branch runs when the LLVM lane count changes (e.g. Float64x2 -> Float32x4: n=2, result size 4). The element type narrows (f64->f32) but the vector lane count increases from 2 to 4, and the shuffle zero-fills the added upper lanes. Calling this a "narrowing conversion" that leaves "unused upper lanes zero" reads as if lanes are dropped. Behavior is correct; the comment wording is misleading.

Comment thread ssa/simd_permute.go

func (b Builder) simdPermute(op SIMDOp, x, indices Expr) Expr {
if op == SIMDLookupOrZero {
name := "llvm.aarch64.neon.tbl1"

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P3] LookupOrZero has no amd64 guard, would emit NEON intrinsic

SIMDLookupOrZero selects llvm.aarch64.neon.tbl1 for everything except wasm. This is correct today only because archsimd exposes LookupOrZero solely on arm64/wasm, but simdOperations registers it with no architecture guard. If the method ever appears on amd64, this would silently emit an AArch64 intrinsic into an amd64 module (invalid IR) rather than failing. Consider an explicit arch check with a panic/unimplemented fallback for the unexpected case.

@github-actions

github-actions Bot commented Oct 3, 2026 •

Copy link
Copy Markdown

LLGo WebAssembly build benchmarks

36caa4db4a8a | workflow run | long-term charts

WebAssembly output sizes
Example, profile and compiler Wasm module vs base Generated JS glue vs base
cprintf/j32-emscripten/LLGo 149240 B 0 B / +0.0% 88157 B 0 B / +0.0%
cprintf/j32-goos-js/LLGo 148511 B 0 B / +0.0% 72985 B 0 B / +0.0%
cprintf/j64-emscripten-memory64/LLGo 136707 B 0 B / +0.0% 91328 B 0 B / +0.0%
cprintf/w32-goos-wasip1/LLGo 153508 B 0 B / +0.0% 0 B 0 B / 0.0%
cprintf/w32-wasi/LLGo 154198 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/j32-emscripten/LLGo 3056999 B +7873 B / +0.3% (worse) 130373 B 0 B / +0.0%
fmtprintf/j32-goos-js/Go 2526852 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/j32-goos-js/LLGo 3152979 B +7827 B / +0.2% (worse) 100773 B 0 B / +0.0%
fmtprintf/j64-emscripten-memory64/LLGo 2819092 B +7250 B / +0.3% (worse) 135853 B 0 B / +0.0%
fmtprintf/w32-goos-wasip1/Go 2500019 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/w32-goos-wasip1/LLGo 2354983 B +5647 B / +0.2% (worse) 0 B 0 B / 0.0%
fmtprintf/w32-wasi/LLGo 2351569 B +5593 B / +0.2% (worse) 0 B 0 B / 0.0%
j32-emscripten/LLGo 148582 B 0 B / +0.0% 88157 B 0 B / +0.0%
j32-goos-js/Go 1895533 B 0 B / +0.0% 0 B 0 B / 0.0%
j32-goos-js/LLGo 148002 B 0 B / +0.0% 72985 B 0 B / +0.0%
j64-emscripten-memory64/LLGo 136047 B 0 B / +0.0% 91328 B 0 B / +0.0%
reflectcall/j32-emscripten/LLGo 1481123 B +7864 B / +0.5% (worse) 104917 B 0 B / +0.0%
reflectcall/j32-goos-js/Go 2191221 B 0 B / +0.0% 0 B 0 B / 0.0%
reflectcall/j32-goos-js/LLGo 1530979 B +7793 B / +0.5% (worse) 89743 B 0 B / +0.0%
reflectcall/j64-emscripten-memory64/LLGo 1373248 B +7223 B / +0.5% (worse) 109492 B 0 B / +0.0%
reflectcall/w32-goos-wasip1/Go 2205707 B 0 B / +0.0% 0 B 0 B / 0.0%
reflectcall/w32-goos-wasip1/LLGo 1283431 B +5655 B / +0.4% (worse) 0 B 0 B / 0.0%
reflectcall/w32-wasi/LLGo 1280963 B +5601 B / +0.4% (worse) 0 B 0 B / 0.0%
w32-goos-wasip1/Go 1909947 B 0 B / +0.0% 0 B 0 B / 0.0%
w32-goos-wasip1/LLGo 153155 B 0 B / +0.0% 0 B 0 B / 0.0%
w32-wasi/LLGo 153845 B 0 B / +0.0% 0 B 0 B / 0.0%
LLGo WebAssembly build measurements
Example and profile Build vs base
j32-emscripten 7.633 s -868 ms / -10.2% (better)
j32-goos-js 7.580 s -173.1 ms / -2.2% (better)
j64-emscripten-memory64 6.857 s -428.9 ms / -5.9% (better)
reflectcall/w32-wasi 25.599 s -540.9 ms / -2.1% (better)
w32-goos-wasip1 4.811 s -276.3 ms / -5.4% (better)
w32-wasi 4.719 s +1.913 ms / +0.04056% (worse)

Compared with ec9c2488bc37 measured in the same runner job.

@github-actions

github-actions Bot commented Oct 3, 2026 •

Copy link
Copy Markdown

LLGo baseline benchmarks

36caa4db4a8a | workflow run | long-term charts

Program measurements

Platform Workload File size vs base Text size vs base Build vs base Run vs base
Linux cprintf 8568 B 0 B / +0.0% 387 B 0 B / +0.0% 1.005 s -151.1 ms / -13.1% (better) 1.217 ms -122.9 us / -9.2% (better)
Linux cprintf-lto 8408 B 0 B / +0.0% 368 B 0 B / +0.0% 945.177 ms -53.36 ms / -5.3% (better) 1.214 ms -77.32 us / -6.0% (better)
Linux fmtprintf 4886552 B +5608 B / +0.1% (worse) 504761 B +3142 B / +0.6% (worse) 7.138 s -471.3 ms / -6.2% (better) 3.047 ms -57.44 us / -1.9% (better)
Linux fmtprintf-lto 3616576 B +15768 B / +0.4% (worse) 443075 B +4537 B / +1.0% (worse) 16.471 s -57.55 ms / -0.3% (better) 3.072 ms +184.6 us / +6.4% (worse)
Linux println 675280 B +328 B / +0.0486% (worse) 16855 B 0 B / +0.0% 1.001 s -13.9 ms / -1.4% (better) 1.591 ms +669 ns / +0.04206% (worse)
Linux println-lto 190040 B +16 B / +0.00842% (worse) 14273 B 0 B / +0.0% 1.262 s -62.17 ms / -4.7% (better) 1.603 ms -24.98 us / -1.5% (better)
macOS cprintf 50736 B 0 B / +0.0% 4409 B 0 B / +0.0% 1.342 s -93.75 ms / -6.5% (better) 3.979 ms -2.067 ms / -34.2% (better)
macOS cprintf-lto 50496 B 0 B / +0.0% 161 B 0 B / +0.0% 1.193 s -673.4 ms / -36.1% (better) 3.262 ms -3.655 ms / -52.8% (better)
macOS fmtprintf 1774368 B +912 B / +0.1% (worse) 880529 B +2749 B / +0.3% (worse) 6.296 s +521.7 ms / +9.0% (worse) 13.866 ms +3.652 ms / +35.8% (worse)
macOS fmtprintf-lto 1361232 B +304 B / +0.02234% (worse) 764957 B +2529 B / +0.3% (worse) 13.564 s +789.2 ms / +6.2% (worse) 4.705 ms -1.091 ms / -18.8% (better)
macOS println 99344 B 0 B / +0.0% 24216 B 0 B / +0.0% 1.186 s -227.6 ms / -16.1% (better) 4.680 ms -2.325 ms / -33.2% (better)
macOS println-lto 83664 B 0 B / +0.0% 21457 B 0 B / +0.0% 1.491 s -95.37 ms / -6.0% (better) 6.515 ms +191.2 us / +3.0% (worse)
Windows MinGW cprintf 651776 B +512 B / +0.1% (worse) 4550 B 0 B / +0.0% 1.309 s +193.9 ms / +17.4% (worse) 2.344 ms +188.9 us / +8.8% (worse)
Windows MinGW cprintf-lto 43520 B 0 B / +0.0% 4486 B 0 B / +0.0% 1.541 s +32.6 ms / +2.2% (worse) 2.237 ms -120 us / -5.1% (better)
Windows MinGW fmtprintf 5440512 B +7680 B / +0.1% (worse) 604454 B +4144 B / +0.7% (worse) 6.074 s +279.6 ms / +4.8% (worse) 5.206 ms -227.5 us / -4.2% (better)
Windows MinGW fmtprintf-lto 4144128 B +16896 B / +0.4% (worse) 551574 B +4992 B / +0.9% (worse) 11.333 s -543.9 ms / -4.6% (better) 5.248 ms +131 us / +2.6% (worse)
Windows MinGW println 707072 B +512 B / +0.1% (worse) 25190 B 0 B / +0.0% 1.391 s +100.4 ms / +7.8% (worse) 4.655 ms +397.9 us / +9.3% (worse)
Windows MinGW println-lto 208896 B 0 B / +0.0% 22054 B 0 B / +0.0% 1.580 s +159.8 ms / +11.3% (worse) 4.226 ms +203.5 us / +5.1% (worse)
Windows MinGW 386 cprintf 601600 B 0 B / +0.0% 5326 B 0 B / +0.0% 1.440 s +174.7 ms / +13.8% (worse) 3.302 ms -146 us / -4.2% (better)
Windows MinGW 386 cprintf-lto 103424 B 0 B / +0.0% 5094 B 0 B / +0.0% 1.454 s +10.23 ms / +0.7% (worse) 3.351 ms -37.1 us / -1.1% (better)
Windows MinGW 386 fmtprintf 4746240 B +1024 B / +0.02158% (worse) 472478 B 0 B / +0.0% 6.131 s -39.89 ms / -0.6% (better) 6.341 ms -648.2 us / -9.3% (better)
Windows MinGW 386 fmtprintf-lto 4148736 B 0 B / +0.0% 451114 B 0 B / +0.0% 11.120 s +83.72 ms / +0.8% (worse) 6.696 ms +46.4 us / +0.7% (worse)
Windows MinGW 386 println 653312 B 0 B / +0.0% 21490 B 0 B / +0.0% 1.468 s +238.5 ms / +19.4% (worse) 5.569 ms -265.8 us / -4.6% (better)
Windows MinGW 386 println-lto 258560 B 0 B / +0.0% 19306 B 0 B / +0.0% 1.703 s +139.3 ms / +8.9% (worse) 5.486 ms -263.9 us / -4.6% (better)
Windows MinGW ARM64 cprintf 661504 B 0 B / +0.0% 4408 B 0 B / +0.0% 1.927 s +46.97 ms / +2.5% (worse) 6.267 ms -176 us / -2.7% (better)
Windows MinGW ARM64 cprintf-lto 43520 B 0 B / +0.0% 4340 B 0 B / +0.0% 1.974 s +78.16 ms / +4.1% (worse) 6.154 ms +14.8 us / +0.2% (worse)
Windows MinGW ARM64 fmtprintf 5348864 B +5120 B / +0.1% (worse) 512636 B +1760 B / +0.3% (worse) 7.122 s +156.3 ms / +2.2% (worse) 12.600 ms +591.1 us / +4.9% (worse)
Windows MinGW ARM64 fmtprintf-lto 4312064 B +9728 B / +0.2% (worse) 478828 B +1564 B / +0.3% (worse) 13.957 s +163.2 ms / +1.2% (worse) 12.748 ms -414.7 us / -3.2% (better)
Windows MinGW ARM64 println 714752 B +512 B / +0.1% (worse) 23884 B 0 B / +0.0% 1.937 s +95.43 ms / +5.2% (worse) 11.002 ms +474.4 us / +4.5% (worse)
Windows MinGW ARM64 println-lto 215552 B 0 B / +0.0% 21232 B 0 B / +0.0% 2.214 s +114.3 ms / +5.4% (worse) 10.912 ms +395.4 us / +3.8% (worse)
Windows MSVC cprintf 893952 B 0 B / +0.0% 65798 B 0 B / +0.0% 1.276 s +16.68 ms / +1.3% (worse) 2.814 ms +52.1 us / +1.9% (worse)
Windows MSVC cprintf-lto 289792 B 0 B / +0.0% 65734 B 0 B / +0.0% 1.451 s +106.7 ms / +7.9% (worse) 2.791 ms +34.7 us / +1.3% (worse)
Windows MSVC fmtprintf 5741056 B +8192 B / +0.1% (worse) 700022 B +4160 B / +0.6% (worse) 5.814 s -10.71 ms / -0.2% (better) 7.244 ms -343.1 us / -4.5% (better)
Windows MSVC fmtprintf-lto 4466688 B +16896 B / +0.4% (worse) 651110 B +4976 B / +0.8% (worse) 11.617 s +321.2 ms / +2.8% (worse) 7.831 ms +249.8 us / +3.3% (worse)
Windows MSVC println 1017856 B +512 B / +0.1% (worse) 120854 B 0 B / +0.0% 1.252 s -28.95 ms / -2.3% (better) 5.540 ms -544.6 us / -9.0% (better)
Windows MSVC println-lto 528384 B 0 B / +0.0% 118390 B 0 B / +0.0% 1.459 s -59.79 ms / -3.9% (better) 5.508 ms +4.5 us / +0.1% (worse)
Windows MSVC 386 cprintf 513536 B 0 B / +0.0% 3931 B 0 B / +0.0% 1.465 s +2.783 ms / +0.2% (worse) 6.701 ms +94.9 us / +1.4% (worse)
Windows MSVC 386 cprintf-lto 44032 B 0 B / +0.0% 3853 B 0 B / +0.0% 1.505 s -87.71 ms / -5.5% (better) 6.936 ms +576.9 us / +9.1% (worse)
Windows MSVC 386 fmtprintf 4480000 B +512 B / +0.01143% (worse) 455868 B 0 B / +0.0% 6.950 s +216.8 ms / +3.2% (worse) 13.166 ms +831.5 us / +6.7% (worse)
Windows MSVC 386 fmtprintf-lto 3899904 B +512 B / +0.01313% (worse) 426651 B 0 B / +0.0% 13.181 s +232.4 ms / +1.8% (worse) 12.460 ms +398.7 us / +3.3% (worse)
Windows MSVC 386 println 567808 B +512 B / +0.1% (worse) 20340 B 0 B / +0.0% 1.471 s +26.76 ms / +1.9% (worse) 10.668 ms +954 us / +9.8% (worse)
Windows MSVC 386 println-lto 199168 B 0 B / +0.0% 18501 B 0 B / +0.0% 1.751 s +25.23 ms / +1.5% (worse) 10.817 ms +256.5 us / +2.4% (worse)
Windows MSVC ARM64 cprintf 662528 B 0 B / +0.0% 4192 B 0 B / +0.0% 1.671 s -13.31 ms / -0.8% (better) 7.494 ms -154.7 us / -2.0% (better)
Windows MSVC ARM64 cprintf-lto 47616 B 0 B / +0.0% 4084 B 0 B / +0.0% 1.663 s -34.55 ms / -2.0% (better) 7.639 ms -303.9 us / -3.8% (better)
Windows MSVC ARM64 fmtprintf 5345280 B +5632 B / +0.1% (worse) 512580 B +1772 B / +0.3% (worse) 6.898 s -10.36 ms / -0.1% (better) 15.514 ms +95.3 us / +0.6% (worse)
Windows MSVC ARM64 fmtprintf-lto 4318720 B +9216 B / +0.2% (worse) 479488 B +1564 B / +0.3% (worse) 13.294 s -112.6 ms / -0.8% (better) 15.380 ms -607.9 us / -3.8% (better)
Windows MSVC ARM64 println 715776 B +512 B / +0.1% (worse) 23908 B 0 B / +0.0% 1.675 s -4.737 ms / -0.3% (better) 12.939 ms -837.1 us / -6.1% (better)
Windows MSVC ARM64 println-lto 220672 B 0 B / +0.0% 21380 B 0 B / +0.0% 1.903 s -16.96 ms / -0.9% (better) 12.996 ms -778.8 us / -5.7% (better)
Core language and compiler benchmarks
Platform Benchmark ns/op vs base
Linux BenchmarkLookupPCRandom 14.430 ns/op -0.19 ns/op / -1.3% (better)
Linux BenchmarkMergeCompilerFlags 191.500 ns/op -8.5 ns/op / -4.2% (better)
Linux BenchmarkMergeLinkerFlags 130.400 ns/op -0.9 ns/op / -0.7% (better)
Linux BenchmarkChannelBuffered 54.940 ns/op -0.09 ns/op / -0.2% (better)
Linux BenchmarkChannelHandoff 13746 ns/op -94 ns/op / -0.7% (better)
Linux BenchmarkDefer 47.450 ns/op -2.57 ns/op / -5.1% (better)
Linux BenchmarkDirectCall 1.515 ns/op +0.349 ns/op / +29.9% (worse)
Linux BenchmarkGlobalRead 1.555 ns/op +0.372 ns/op / +31.4% (worse)
Linux BenchmarkGlobalWrite 7.767 ns/op -0.023 ns/op / -0.3% (better)
Linux BenchmarkGoroutine 23514 ns/op -34 ns/op / -0.1% (better)
Linux BenchmarkInterfaceCall 6.213 ns/op +0.368 ns/op / +6.3% (worse)
Linux BenchmarkRuntimeGetG 2.896 ns/op -0.073 ns/op / -2.5% (better)
macOS BenchmarkLookupPCRandom 18.200 ns/op +5.2 ns/op / +40.0% (worse)
macOS BenchmarkMergeCompilerFlags 169.900 ns/op +54.5 ns/op / +47.2% (worse)
macOS BenchmarkMergeLinkerFlags 103.400 ns/op +24.74 ns/op / +31.5% (worse)
macOS BenchmarkChannelBuffered 37.940 ns/op +5.19 ns/op / +15.8% (worse)
macOS BenchmarkChannelHandoff 7160 ns/op -3996 ns/op / -35.8% (better)
macOS BenchmarkDefer 51.030 ns/op +6.32 ns/op / +14.1% (worse)
macOS BenchmarkDirectCall 1.465 ns/op +0.233 ns/op / +18.9% (worse)
macOS BenchmarkGlobalRead 1.369 ns/op +0.148 ns/op / +12.1% (worse)
macOS BenchmarkGlobalWrite 1.799 ns/op +0.552 ns/op / +44.3% (worse)
macOS BenchmarkGoroutine 46966 ns/op -10115 ns/op / -17.7% (better)
macOS BenchmarkInterfaceCall 4.971 ns/op +0.847 ns/op / +20.5% (worse)
macOS BenchmarkRuntimeGetG 3.291 ns/op +0.596 ns/op / +22.1% (worse)
Windows MinGW BenchmarkLookupPCRandom 6.990 ns/op -0.094 ns/op / -1.3% (better)
Windows MinGW BenchmarkMergeCompilerFlags 724.600 ns/op +146.2 ns/op / +25.3% (worse)
Windows MinGW BenchmarkMergeLinkerFlags 544.300 ns/op +8.6 ns/op / +1.6% (worse)
Windows MinGW BenchmarkChannelBuffered 28.370 ns/op +0.25 ns/op / +0.9% (worse)
Windows MinGW BenchmarkChannelHandoff 796.200 ns/op +2.1 ns/op / +0.3% (worse)
Windows MinGW BenchmarkDefer 35.960 ns/op +2.56 ns/op / +7.7% (worse)
Windows MinGW BenchmarkDirectCall 0.905 ns/op +0.0096 ns/op / +1.1% (worse)
Windows MinGW BenchmarkGlobalRead 0.915 ns/op -0.0072 ns/op / -0.8% (better)
Windows MinGW BenchmarkGlobalWrite 4.531 ns/op +0.008 ns/op / +0.2% (worse)
Windows MinGW BenchmarkGoroutine 42376 ns/op +1528 ns/op / +3.7% (worse)
Windows MinGW BenchmarkInterfaceCall 4.482 ns/op +0.017 ns/op / +0.4% (worse)
Windows MinGW BenchmarkRuntimeGetG 0.919 ns/op -0.1315 ns/op / -12.5% (better)
Windows MinGW 386 BenchmarkLookupPCRandom 15.680 ns/op +0.07 ns/op / +0.4% (worse)
Windows MinGW 386 BenchmarkMergeCompilerFlags 975.600 ns/op -79.4 ns/op / -7.5% (better)
Windows MinGW 386 BenchmarkMergeLinkerFlags 1057 ns/op -132 ns/op / -11.1% (better)
Windows MinGW 386 BenchmarkChannelBuffered 36.450 ns/op -0.25 ns/op / -0.7% (better)
Windows MinGW 386 BenchmarkChannelHandoff 624.800 ns/op -7.8 ns/op / -1.2% (better)
Windows MinGW 386 BenchmarkDefer 25.900 ns/op -1.77 ns/op / -6.4% (better)
Windows MinGW 386 BenchmarkDirectCall 0.888 ns/op -0.0523 ns/op / -5.6% (better)
Windows MinGW 386 BenchmarkGlobalRead 0.896 ns/op -0.0426 ns/op / -4.5% (better)
Windows MinGW 386 BenchmarkGlobalWrite 8.091 ns/op -0.324 ns/op / -3.9% (better)
Windows MinGW 386 BenchmarkGoroutine 49223 ns/op -3628 ns/op / -6.9% (better)
Windows MinGW 386 BenchmarkInterfaceCall 4.479 ns/op -0.509 ns/op / -10.2% (better)
Windows MinGW 386 BenchmarkRuntimeGetG 0.952 ns/op -0.0356 ns/op / -3.6% (better)
Windows MinGW ARM64 BenchmarkLookupPCRandom 12.120 ns/op +0.01 ns/op / +0.1% (worse)
Windows MinGW ARM64 BenchmarkMergeCompilerFlags 583.800 ns/op +0.5 ns/op / +0.1% (worse)
Windows MinGW ARM64 BenchmarkMergeLinkerFlags 542.400 ns/op -3.4 ns/op / -0.6% (better)
Windows MinGW ARM64 BenchmarkChannelBuffered 39.010 ns/op -0.47 ns/op / -1.2% (better)
Windows MinGW ARM64 BenchmarkChannelHandoff 3555 ns/op +1163 ns/op / +48.6% (worse)
Windows MinGW ARM64 BenchmarkDefer 57.930 ns/op +1.51 ns/op / +2.7% (worse)
Windows MinGW ARM64 BenchmarkDirectCall 0.663 ns/op +0.0001 ns/op / +0.01508% (worse)
Windows MinGW ARM64 BenchmarkGlobalRead 0.664 ns/op +0.0001 ns/op / +0.01507% (worse)
Windows MinGW ARM64 BenchmarkGlobalWrite 0.737 ns/op +0.0002 ns/op / +0.02714% (worse)
Windows MinGW ARM64 BenchmarkGoroutine 65555 ns/op +2209 ns/op / +3.5% (worse)
Windows MinGW ARM64 BenchmarkInterfaceCall 4.239 ns/op +0.093 ns/op / +2.2% (worse)
Windows MinGW ARM64 BenchmarkRuntimeGetG 1.769 ns/op 0 ns/op / +0.0%
Windows MSVC BenchmarkLookupPCRandom 9.671 ns/op -0.249 ns/op / -2.5% (better)
Windows MSVC BenchmarkMergeCompilerFlags 403.700 ns/op +21.3 ns/op / +5.6% (worse)
Windows MSVC BenchmarkMergeLinkerFlags 338.700 ns/op -4.1 ns/op / -1.2% (better)
Windows MSVC BenchmarkChannelBuffered 24.850 ns/op +0.65 ns/op / +2.7% (worse)
Windows MSVC BenchmarkChannelHandoff 967.700 ns/op -52.3 ns/op / -5.1% (better)
Windows MSVC BenchmarkDefer 44.420 ns/op -0.82 ns/op / -1.8% (better)
Windows MSVC BenchmarkDirectCall 1.356 ns/op 0 ns/op / +0.0%
Windows MSVC BenchmarkGlobalRead 1.356 ns/op -0.001 ns/op / -0.1% (better)
Windows MSVC BenchmarkGlobalWrite 2.166 ns/op +0.001 ns/op / +0.04619% (worse)
Windows MSVC BenchmarkGoroutine 58101 ns/op +1056 ns/op / +1.9% (worse)
Windows MSVC BenchmarkInterfaceCall 7.105 ns/op -0.307 ns/op / -4.1% (better)
Windows MSVC BenchmarkRuntimeGetG 1.902 ns/op -0.001 ns/op / -0.1% (better)
Windows MSVC 386 BenchmarkLookupPCRandom 70.990 ns/op +1.62 ns/op / +2.3% (worse)
Windows MSVC 386 BenchmarkMergeCompilerFlags 765.100 ns/op +7.8 ns/op / +1.0% (worse)
Windows MSVC 386 BenchmarkMergeLinkerFlags 709.800 ns/op +3.4 ns/op / +0.5% (worse)
Windows MSVC 386 BenchmarkChannelBuffered 54.880 ns/op +0.52 ns/op / +1.0% (worse)
Windows MSVC 386 BenchmarkChannelHandoff 1425 ns/op -373 ns/op / -20.7% (better)
Windows MSVC 386 BenchmarkDefer 48.420 ns/op +3.82 ns/op / +8.6% (worse)
Windows MSVC 386 BenchmarkDirectCall 1.102 ns/op -0.004 ns/op / -0.4% (better)
Windows MSVC 386 BenchmarkGlobalRead 1.113 ns/op +0.033 ns/op / +3.1% (worse)
Windows MSVC 386 BenchmarkGlobalWrite 17.640 ns/op +0.58 ns/op / +3.4% (worse)
Windows MSVC 386 BenchmarkGoroutine 250534 ns/op -9963 ns/op / -3.8% (better)
Windows MSVC 386 BenchmarkInterfaceCall 5.807 ns/op +0.135 ns/op / +2.4% (worse)
Windows MSVC 386 BenchmarkRuntimeGetG 1.667 ns/op +0.054 ns/op / +3.3% (worse)
Windows MSVC ARM64 BenchmarkLookupPCRandom 12.140 ns/op 0 ns/op / +0.0%
Windows MSVC ARM64 BenchmarkMergeCompilerFlags 587 ns/op -3.4 ns/op / -0.6% (better)
Windows MSVC ARM64 BenchmarkMergeLinkerFlags 533.700 ns/op -9 ns/op / -1.7% (better)
Windows MSVC ARM64 BenchmarkChannelBuffered 37.470 ns/op -0.09 ns/op / -0.2% (better)
Windows MSVC ARM64 BenchmarkChannelHandoff 3013 ns/op -199 ns/op / -6.2% (better)
Windows MSVC ARM64 BenchmarkDefer 62.880 ns/op +0.56 ns/op / +0.9% (worse)
Windows MSVC ARM64 BenchmarkDirectCall 0.663 ns/op -0.0001 ns/op / -0.01507% (better)
Windows MSVC ARM64 BenchmarkGlobalRead 0.664 ns/op +0.0001 ns/op / +0.01507% (worse)
Windows MSVC ARM64 BenchmarkGlobalWrite 3.797 ns/op +0.001 ns/op / +0.02634% (worse)
Windows MSVC ARM64 BenchmarkGoroutine 63437 ns/op +1663 ns/op / +2.7% (worse)
Windows MSVC ARM64 BenchmarkInterfaceCall 4.231 ns/op +0.088 ns/op / +2.1% (worse)
Windows MSVC ARM64 BenchmarkRuntimeGetG 1.809 ns/op +0.04 ns/op / +2.3% (worse)
Timer runtime benchmarks
Platform Operation and runtime ns/op vs base
Linux AfterFuncZeroDelivery/Go 902.400 ns/op +1.4 ns/op / +0.2% (worse)
Linux AfterFuncZeroDelivery/LLGo 45532 ns/op +8886 ns/op / +24.2% (worse)
Linux CreateStop/Go 288.400 ns/op -4.4 ns/op / -1.5% (better)
Linux CreateStop/LLGo 1716 ns/op -271 ns/op / -13.6% (better)
Linux RearmStopped/Go 114.900 ns/op +0.3 ns/op / +0.3% (worse)
Linux RearmStopped/LLGo 1335 ns/op -127 ns/op / -8.7% (better)
Linux ResetActive/Go 67.420 ns/op -0.17 ns/op / -0.3% (better)
Linux ResetActive/LLGo 705.700 ns/op -164.8 ns/op / -18.9% (better)
Linux ResetHeap1024/Go 67.330 ns/op -0.09 ns/op / -0.1% (better)
Linux ResetHeap1024/LLGo 176.600 ns/op +3.2 ns/op / +1.8% (worse)
macOS AfterFuncZeroDelivery/Go 470.500 ns/op -89.3 ns/op / -16.0% (better)
macOS AfterFuncZeroDelivery/LLGo 95949 ns/op -2835 ns/op / -2.9% (better)
macOS CreateStop/Go 207.700 ns/op +69.1 ns/op / +49.9% (worse)
macOS CreateStop/LLGo 571.100 ns/op -31 ns/op / -5.1% (better)
macOS RearmStopped/Go 73.870 ns/op +10.85 ns/op / +17.2% (worse)
macOS RearmStopped/LLGo 518 ns/op -34.7 ns/op / -6.3% (better)
macOS ResetActive/Go 54.150 ns/op +10.79 ns/op / +24.9% (worse)
macOS ResetActive/LLGo 225.400 ns/op -29.3 ns/op / -11.5% (better)
macOS ResetHeap1024/Go 55.700 ns/op +9.3 ns/op / +20.0% (worse)
macOS ResetHeap1024/LLGo 163.800 ns/op +60.1 ns/op / +58.0% (worse)
Windows MinGW AfterFuncZeroDelivery/Go 407 ns/op -27.1 ns/op / -6.2% (better)
Windows MinGW AfterFuncZeroDelivery/LLGo 88801 ns/op -33 ns/op / -0.03715% (better)
Windows MinGW CreateStop/Go 118.900 ns/op -0.8 ns/op / -0.7% (better)
Windows MinGW CreateStop/LLGo 306.100 ns/op +0.9 ns/op / +0.3% (worse)
Windows MinGW RearmStopped/Go 46.180 ns/op -1.83 ns/op / -3.8% (better)
Windows MinGW RearmStopped/LLGo 185.600 ns/op -3.3 ns/op / -1.7% (better)
Windows MinGW ResetActive/Go 20.780 ns/op -0.07 ns/op / -0.3% (better)
Windows MinGW ResetActive/LLGo 449.900 ns/op -7.2 ns/op / -1.6% (better)
Windows MinGW ResetHeap1024/Go 20.460 ns/op -0.63 ns/op / -3.0% (better)
Windows MinGW ResetHeap1024/LLGo 89.210 ns/op -1.62 ns/op / -1.8% (better)
Windows MinGW 386 AfterFuncZeroDelivery/Go 607.200 ns/op -6.8 ns/op / -1.1% (better)
Windows MinGW 386 AfterFuncZeroDelivery/LLGo 97466 ns/op -5238 ns/op / -5.1% (better)
Windows MinGW 386 CreateStop/Go 159.300 ns/op -2.9 ns/op / -1.8% (better)
Windows MinGW 386 CreateStop/LLGo 385.300 ns/op +16.5 ns/op / +4.5% (worse)
Windows MinGW 386 RearmStopped/Go 59.560 ns/op -2.07 ns/op / -3.4% (better)
Windows MinGW 386 RearmStopped/LLGo 240.600 ns/op -2.8 ns/op / -1.2% (better)
Windows MinGW 386 ResetActive/Go 29.050 ns/op -0.25 ns/op / -0.9% (better)
Windows MinGW 386 ResetActive/LLGo 630.100 ns/op -56.3 ns/op / -8.2% (better)
Windows MinGW 386 ResetHeap1024/Go 28.140 ns/op -1.76 ns/op / -5.9% (better)
Windows MinGW 386 ResetHeap1024/LLGo 116.900 ns/op -3.9 ns/op / -3.2% (better)
Windows MinGW ARM64 AfterFuncZeroDelivery/Go 671.100 ns/op +5.5 ns/op / +0.8% (worse)
Windows MinGW ARM64 AfterFuncZeroDelivery/LLGo 149275 ns/op +3400 ns/op / +2.3% (worse)
Windows MinGW ARM64 CreateStop/Go 204.200 ns/op +0.2 ns/op / +0.1% (worse)
Windows MinGW ARM64 CreateStop/LLGo 354.300 ns/op -9.5 ns/op / -2.6% (better)
Windows MinGW ARM64 RearmStopped/Go 70.550 ns/op -0.08 ns/op / -0.1% (better)
Windows MinGW ARM64 RearmStopped/LLGo 248.500 ns/op -1.2 ns/op / -0.5% (better)
Windows MinGW ARM64 ResetActive/Go 31.090 ns/op +0.06 ns/op / +0.2% (worse)
Windows MinGW ARM64 ResetActive/LLGo 122.900 ns/op +0.7 ns/op / +0.6% (worse)
Windows MinGW ARM64 ResetHeap1024/Go 31.030 ns/op -0.04 ns/op / -0.1% (better)
Windows MinGW ARM64 ResetHeap1024/LLGo 125.600 ns/op +0.9 ns/op / +0.7% (worse)
Windows MSVC AfterFuncZeroDelivery/Go 396.400 ns/op +17.5 ns/op / +4.6% (worse)
Windows MSVC AfterFuncZeroDelivery/LLGo 111453 ns/op +3696 ns/op / +3.4% (worse)
Windows MSVC CreateStop/Go 89.890 ns/op -0.15 ns/op / -0.2% (better)
Windows MSVC CreateStop/LLGo 352.100 ns/op +1.3 ns/op / +0.4% (worse)
Windows MSVC RearmStopped/Go 24.480 ns/op +0.03 ns/op / +0.1% (worse)
Windows MSVC RearmStopped/LLGo 218 ns/op +0.8 ns/op / +0.4% (worse)
Windows MSVC ResetActive/Go 14.760 ns/op -0.03 ns/op / -0.2% (better)
Windows MSVC ResetActive/LLGo 114 ns/op -1.2 ns/op / -1.0% (better)
Windows MSVC ResetHeap1024/Go 14.770 ns/op -0.17 ns/op / -1.1% (better)
Windows MSVC ResetHeap1024/LLGo 105.300 ns/op +0.8 ns/op / +0.8% (worse)
Windows MSVC 386 AfterFuncZeroDelivery/Go 1032 ns/op +32 ns/op / +3.2% (worse)
Windows MSVC 386 AfterFuncZeroDelivery/LLGo 363408 ns/op +10689 ns/op / +3.0% (worse)
Windows MSVC 386 CreateStop/Go 261.100 ns/op +3.9 ns/op / +1.5% (worse)
Windows MSVC 386 CreateStop/LLGo 724.500 ns/op +32.1 ns/op / +4.6% (worse)
Windows MSVC 386 RearmStopped/Go 95.100 ns/op +2.46 ns/op / +2.7% (worse)
Windows MSVC 386 RearmStopped/LLGo 462.700 ns/op -10.7 ns/op / -2.3% (better)
Windows MSVC 386 ResetActive/Go 45.940 ns/op +1.3 ns/op / +2.9% (worse)
Windows MSVC 386 ResetActive/LLGo 297.800 ns/op -43.8 ns/op / -12.8% (better)
Windows MSVC 386 ResetHeap1024/Go 46.040 ns/op +0.77 ns/op / +1.7% (worse)
Windows MSVC 386 ResetHeap1024/LLGo 173.100 ns/op +7.2 ns/op / +4.3% (worse)
Windows MSVC ARM64 AfterFuncZeroDelivery/Go 675.600 ns/op -4.4 ns/op / -0.6% (better)
Windows MSVC ARM64 AfterFuncZeroDelivery/LLGo 179424 ns/op -59 ns/op / -0.03287% (better)
Windows MSVC ARM64 CreateStop/Go 200.400 ns/op +1.2 ns/op / +0.6% (worse)
Windows MSVC ARM64 CreateStop/LLGo 395.200 ns/op -6.7 ns/op / -1.7% (better)
Windows MSVC ARM64 RearmStopped/Go 70.510 ns/op -0.07 ns/op / -0.1% (better)
Windows MSVC ARM64 RearmStopped/LLGo 273.700 ns/op -0.8 ns/op / -0.3% (better)
Windows MSVC ARM64 ResetActive/Go 31.060 ns/op -0.07 ns/op / -0.2% (better)
Windows MSVC ARM64 ResetActive/LLGo 135 ns/op +2.8 ns/op / +2.1% (worse)
Windows MSVC ARM64 ResetHeap1024/Go 31.110 ns/op -0.07 ns/op / -0.2% (better)
Windows MSVC ARM64 ResetHeap1024/LLGo 137.700 ns/op +0.3 ns/op / +0.2% (worse)

Compared with ec9c2488bc37 measured in the same runner job.

@zhouguangyuan0718

Copy link
Copy Markdown
Collaborator Author

@fennoai 请重新检视这个 PR 的最新提交 a2b7c8ba703f10a548e10e4e6bd268f959137149,并以当前主干 ec9c2488bc3730dcb5f017c4319b48b328e2ebd6 为参照。

这个 PR 创建较早,请特别检查哪些兼容或补救逻辑已经被近期主干改动覆盖,是否可以删除或收窄。重点包括 Emscripten SIMD 调用桥接与新的异常处理/线程模式、WASI LTO、CPU 初始化,以及源码补丁与测试集成。也请复查 SIMD lowering 的正确性及之前的检视意见,给出具体文件、触发条件和验证依据。

@fennoai

fennoai Bot commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

An internal error occurred. Please try again later.

View job run

@zhouguangyuan0718

Copy link
Copy Markdown
Collaborator Author

@fennoai review

Please retry the review after the internal error. Review PR head a2b7c8b against current main ec9c248, focusing on correctness and compatibility code that is now redundant or too broad after recent mainline changes.

@fennoai

fennoai Bot commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

An internal error occurred. Please try again later.

View job run

@zhouguangyuan0718

Copy link
Copy Markdown
Collaborator Author

@fennoai review

Please review the current PR head 1f91d8ee519848f564c2805f4b77be2e4d9843e4 against base ec9c2488bc3730dcb5f017c4319b48b328e2ebd6. Focus on correctness and concrete simplifications: redundant compatibility/workaround logic, duplicated SIMD lowering or flag parsing, and code made unnecessary by recent mainline changes. Preserve the existing js/wasm exception mechanism. Please distinguish actionable bugs from optional refactoring, and give exact files, triggering cases, and supporting evidence. Earlier retry requests failed with internal errors; please run a fresh review of this head.

@fennoai

fennoai Bot commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

Review: PR #2722 @ 1f91d8ee vs main ec9c2488

The head is rebased directly onto the current main tip (merge-base = ec9c2488), so the range below is PR-only. I re-examined the compatibility/workaround logic against recent mainline changes (native EH, pthread features, WASI LTO, CPU init, source patches) and re-verified SIMD lowering correctness. No actionable bugs found. Several earlier review points are already resolved in this head.

Central question — is the compat logic now redundant or too broad?

The Emscripten SIMD-call bridge is load-bearing, not redundant or over-broad. I traced the default Emscripten configuration: internal/crosscompile/crosscompile.go applies -fwasm-exceptions + native SjLj (-wasm-enable-sjlj, -wasm-use-legacy-eh=false, -exception-model=wasm) only in case "wasip1": (lines 788–807). The Emscripten path case "js": (lines 841–889) instead uses -sASYNCIFY=1 with Emscripten's default JavaScript SjLj — which cannot carry v128.

  • Evidence: needsEmscriptenSIMDCallBridge (internal/build/wasm_simd_calls.go:13) therefore returns true by default for Emscripten/GoJS (no -fwasm-exceptions, SUPPORT_LONGJMP defaults to "1", wasmEH=false ⇒ emscriptenNativeSjLj false). Confirmed by wasm_simd_calls_test.go ({name:"default", bridge:true}, default flags ⇒ 4 bridges).
  • It is correctly narrowed: when a user opts into native Wasm SjLj (-fwasm-exceptions / -sSUPPORT_LONGJMP=wasm) via EMCC_CFLAGS, the gate turns the bridge off. The c6f35c796 commit added the GoJS-vs-native-SjLj distinction the earlier feedback asked for.
  • No duplicated flag parser: internal/build/wasm_eh.go (applyEmscriptenEHFeature) only sets the exception-handling/atomics/bulk-memory target-feature attributes from WasmProfile/-pthread; it does not parse emcc SjLj flags. The two answer different questions and are not inconsistent.

Other flagged areas are current, not stale:

  • WASI LTO: -lto=thin/full -target wasi is wired through crosscompile.go:788–801 and documented commands match CI (.github/workflows/llgo.yml:937–940).
  • CPU init (runtime/_patch/internal/cpu/*, runtime/internal/runtime/cpu_env.go, runtime/build.go adding internal/cpu to sourcePatchPkgs): the Windows bridge correctly moved out of link_windows_llgo.go into the patched internal/cpu package; the darwin sysctl shim replaces the deleted os_darwin.go with equivalent behavior. Passing the full GODEBUG string to cpu.Initialize is correct (processOptions ignores non-cpu. fields).
  • Source-patch / test integration (internal/packages/load.go): switching the dedup key from PkgPath to pkg.ID is consistent with the sibling cache store/load (also keyed on pkg.ID) and correctly lets the test-augmented variant receive its own patch files.
  • GOAMD64 (internal/plan9asm/translate.go:106): injecting a single #define GOAMD64_vN matches cmd/go exactly (gc.go:352 emits one -D GOAMD64_<value>, non-cumulative); the nested #ifdef in getGOAMD64level is satisfied.

Correctness — verified sound

  • Shift clamping (ssa/simd_integer.go): counts are clamped >= width → width-1 before any shl/lshr/ashr, and logical shifts additionally select(large, 0, …), so no lane produces LLVM poison; the ARM64 SIMDShift low-byte sign-extension path is correct.
  • Permute index masking (ssa/simd_permute.go): every index is masked & (n-1) before ExtractElement (lane counts are powers of two), so no OOB extract; PermuteOrZero zeroes negative indices.
  • Conversions (ssa/simd_convert.go): fptosi.sat/fptoui.sat avoid NaN/overflow poison; the amd64 sentinel override matches CVTTPS2* "integer indefinite"/all-ones semantics; widening-lane shuffle zero-fills added lanes.
  • Load/store (ssa/simd.go): AssertNilDeref + element alignment, with the *[N]elem pointer type guaranteed by cl/simd.go matches via types.Identical.
  • Emscripten bridge GC-safety: allocate(...) is only reached for VectorTypeKind results/operands (wasm_simd_calls.go), so the untracked stack slots hold only vector bits, never Go pointers — the comment's claim holds.

Earlier-round feedback already addressed

  • simdFeatures dead wrapper removed; SIMD calls requireSIMDFeatures() directly.
  • LookupOrZero now has an explicit arch guard (cl/simd.go:171), so it can no longer emit llvm.aarch64.neon.tbl1 into a non-arm64/wasm module.
  • Comparison predicate tables hoisted to package-level var maps (simdFloatPredicates/simdIntPredicates/simdIntegerIntrinsics) — no per-call allocation.
  • The "narrowing conversions" comment was reworded (simd_convert.go:62).
  • README's prior WASI "blocked at module loading / WAMR" contradiction is resolved — it now documents Wasmer 7.5.0 (dev/install_wasmer.sh exists) and the commands match CI.

Optional (non-blocking) — refactoring, not bugs

  1. cl/simd.go:246 — the 4-conjunct simdConvert guard (d.signature == simdConvert && shape.Len() != lanes.Len() && …IsFloat == 0 && …IsFloat == 0) repeats the .Underlying().(*types.Basic).Info()&types.IsFloat extraction twice on one line. Logic is correct for the registered conversions; extracting fromFloat/toFloat locals would make the "reject lane-count change unless a side is float" rule self-evident.
  2. ssa/simd_integer.go (simdShiftValue) — the arithmetic-right (AShr) path returns without the select(large, 0, …) used by the Shl/LShr paths. This is correct (the top-of-function clamp to width-1 yields sign-extension on over-shift), but a one-line comment on the asymmetry would reduce future-edit fragility.
  3. ssa/simd_permute.go — dynamic (non-constant) amd64 Permute/PermuteOrZero lower to an O(n) per-lane extract/insert gather (baseline SSE2 has no variable byte shuffle). The in-code comment correctly scopes the shufflevector fold to constant indices, so this is an intentional, documented baseline tradeoff — noted only so it is a conscious choice.

Verified with the fixed range git diff ec9c2488bc3730dcb5f017c4319b48b328e2ebd6...1f91d8ee519848f564c2805f4b77be2e4d9843e4. A full go build/go test could not run in this environment (LLVM C headers absent), so the above is static analysis plus cross-checking against cmd/go, internal/cpu, and the PR's own tests.

@xushiwei
xushiwei merged commit f167e3c into xgo-dev:main Oct 6, 2026
92 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants