Skip to content

wasm: enable browser LTO and qualify JS SIMD profiles - #2744

Merged
xushiwei merged 2 commits into
xgo-dev:mainfrom
zhouguangyuan0718:codex/js-simd-lto-20261006
Oct 7, 2026
Merged

xushiwei merged 2 commits into
xgo-dev:mainfrom
zhouguangyuan0718:codex/js-simd-lto-20261006

Conversation

@zhouguangyuan0718

@zhouguangyuan0718 zhouguangyuan0718 commented Oct 6, 2026 •

Copy link
Copy Markdown
Collaborator

Browser builds accepted -lto=thin/full without forwarding -flto to compilation or final linking. Forward both the LTO mode and the requested LTO optimizer level for GoJS, Emscripten, and Emscripten Memory64. Emscripten 6.0.8 consumes -O for compilation and post-link optimization but leaves wasm-ld's LTO level at its default unless --lto-O is supplied explicitly. O0–O3 now select the corresponding LTO level; Os/Oz use LTO O2 while retaining size attributes and post-link size optimization.

Prefer the SDK's emar for package archives and MRI merges, preserving the LLGO_AR override. Resolve Windows SDK .bat tools with LookPath; tests cover a competing SDK on PATH as well as fallback and override behavior.

LLVM 22's O3 argument promotion can also replace an internal SIMD bridge's pointer argument with v128, despite noinline/optnone. That recreates an unsupported JavaScript SjLj signature on Memory64. Make the bridge's vector argument loads volatile to preserve its memory ABI. A regression test reproduces the failure under argument promotion and both O3 pre-link LTO pipelines. The default JavaScript SjLj/Asyncify configuration remains in use.

Add dev/test_wasm_simd.sh to the wasm test-command CI job. Each of the three JS profiles runs an O0 boundary executable, the complete O2 SIMD suite with LTO off/Thin/Full, and O3 Thin/Full LTO boundary executables. The small O0 executable avoids the full test binary's engine local-variable limit.

Validation with Go 1.27.0, LLVM 22.1.8, Emscripten 6.0.8, and Node 24.19.0:

  • All 54 compile/link configuration cases and the three argument-promotion regression cases pass; both regressions fail before their respective fixes. All crosscompile tests and the affected build/archive/bridge tests pass.
  • All 36 boundary combinations (three profiles × six optimization levels × Thin/Full LTO) pass. Verbose SDK traces confirm the actual wasm-ld --lto-O level in every case.
  • All 18 CI-script combinations pass, including nine complete O2 SIMD runs and TestLaneBounds in each, requalifying the three targets from wasm: experimental SIMD vector calls fail at Emscripten JS exception boundaries #2729 after simd: expand SIMD128 operations and runtime integration #2722.
  • Memory64 O3 ThinLTO changes from two v128 JavaScript imports to zero and passes in real Chrome. Full LTO browser boundaries for all three profiles and Memory64 Go/C++ Thin/Full LTO boundaries also passed during initial qualification.
  • Extracted SIMD archive members are LLVM bitcode; compile and final-link traces contain -flto. ShellCheck, shell syntax, and diff checks pass.

The expanded initial internal/build run reached the default 10-minute package timeout after 538 completed top-level tests without earlier top-level failures. The remaining 143 tests passed separately in 55.5 seconds. Affected tests were rerun after these changes; The previous head also passed the full internal/build CI run and both Windows Go jobs, including the new .bat fixture.

Rebased onto main 1866f43f5. The affected compile/link, archive, and SIMD bridge tests pass after rebase. Main already contains the GOROOT deadline-test scheduling fix; its focused regression tests pass, and this PR no longer changes that test.

Refs #2729.

@fennoai fennoai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review summary

This PR extends the SIMD CI matrix to JavaScript/WASM targets (GoJS, Emscripten, Emscripten Memory64) across off/thin/full LTO, adds an emscriptenArchiver() helper so bitcode is archived with the SDK's emar, and forwards -flto=<mode> to both compile and link for the Emscripten branch. The change is well-factored and the docs/comments are accurate.

Verification against the code:

  • dev/test_wasm_simd.sh matches the README description (3 profiles x O0 boundary + O2 off/thin/full, -pclntab=none), and sets GOEXPERIMENT=simd internally.
  • The Emscripten LTO branch (crosscompile.go) correctly appends ltoMode.ClangFlag() to both CCFLAGS and LDFLAGS.
  • emscriptenArchiver() is only consulted for the GoJS/Emscripten providers and keeps the LLGO_AR override first.

No security concerns found. A few minor, non-blocking notes are inline.

Comment thread internal/build/build.go
return ""
}
if cc, err := exec.LookPath(c.crossCompile.CC); err == nil {
if ar, err := exec.LookPath(filepath.Join(filepath.Dir(cc), "emar")); err == nil {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Minor: this sibling-tool probe diverges from the existing siblingTool() helper used just below for llvm-ar. siblingTool handles the Windows .exe suffix explicitly via os.Stat, whereas exec.LookPath of an absolute path relies on %PATHEXT% semantics. The Emscripten SDK ships emar as emar.bat on Windows, so the two lookups can behave differently and the sibling-emar probe may fall through to the bare emar on PATH (which could be a different SDK than the resolved emcc). Consider reusing siblingTool(cc, "emar") after resolving cc, so Windows handling and the lookup strategy stay aligned with the neighboring llvm-ar logic.

Comment thread internal/build/build.go Outdated
}

// Emscripten C++ bitcode can be newer than LLGo's linked LLVM. Its emar
// wrapper selects the SDK's matching llvm-ar for both objects and MRI merges.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Minor (doc clarity): the comment mentions only "Emscripten C++ bitcode", but emscriptenArchiver() also returns emar for WasmProviderGoJS. That is intentional (the GOOS=js path configures CC=emcc and defaults to WasmProviderGoJS), but a one-line note that GoJS also builds through emcc/emar would prevent a future "is the GoJS case a mistake?" question.

}...)
if ltoMode.Enabled() {
export.CCFLAGS = append(export.CCFLAGS, ltoMode.ClangFlag())
export.LDFLAGS = append(export.LDFLAGS, ltoMode.ClangFlag())

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Minor (consistency): this forwards only the base ClangFlag() (e.g. -flto=thin). The sibling WASI branch additionally emits the LTO linker opt level (-Wl,--lto-O2 via ltoLinkerOptFlag) and exception/feature defaults. If emcc is expected to inject the LTO backend opt-level and feature wiring itself, a brief comment stating that would clarify why this branch intentionally differs from the WASI branch's treatment.

Comment thread internal/build/build.go
if ar := os.Getenv("LLGO_AR"); ar != "" {
return ar
}
if ar := c.emscriptenArchiver(); ar != "" {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Minor (perf, pre-existing pattern amplified): archiver() / archiveMerger() are called once per package in createArchiveFile, and on wasm/Emscripten builds emscriptenArchiver() now adds up to ~3 exec.LookPath PATH scans per package (CC, sibling emar, bare emar) on top of the existing llvm-ar scan. The result is deterministic for the lifetime of the context. Consider resolving the archiver once and caching it (guarded with sync.Once, mirroring the existing plan9asmOnce) to avoid hundreds of redundant PATH scans on large builds. Not blocking.

@codecov

codecov Bot commented Oct 6, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@github-actions

github-actions Bot commented Oct 6, 2026 •

Copy link
Copy Markdown

LLGo WebAssembly build benchmarks

6cb2cbe98e0c | workflow run | long-term charts

WebAssembly output sizes
Example, profile and compiler Wasm module vs base Generated JS glue vs base
cprintf/j32-emscripten/LLGo 149240 B 0 B / +0.0% 88157 B 0 B / +0.0%
cprintf/j32-goos-js/LLGo 148511 B 0 B / +0.0% 72985 B 0 B / +0.0%
cprintf/j64-emscripten-memory64/LLGo 136707 B 0 B / +0.0% 91328 B 0 B / +0.0%
cprintf/w32-goos-wasip1/LLGo 154296 B 0 B / +0.0% 0 B 0 B / 0.0%
cprintf/w32-wasi/LLGo 154979 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/j32-emscripten/LLGo 3056999 B 0 B / +0.0% 130373 B 0 B / +0.0%
fmtprintf/j32-goos-js/Go 2526852 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/j32-goos-js/LLGo 3152979 B 0 B / +0.0% 100773 B 0 B / +0.0%
fmtprintf/j64-emscripten-memory64/LLGo 2819092 B 0 B / +0.0% 135853 B 0 B / +0.0%
fmtprintf/w32-goos-wasip1/Go 2500019 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/w32-goos-wasip1/LLGo 2355448 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/w32-wasi/LLGo 2352026 B 0 B / +0.0% 0 B 0 B / 0.0%
j32-emscripten/LLGo 148582 B 0 B / +0.0% 88157 B 0 B / +0.0%
j32-goos-js/Go 1895533 B 0 B / +0.0% 0 B 0 B / 0.0%
j32-goos-js/LLGo 148002 B 0 B / +0.0% 72985 B 0 B / +0.0%
j64-emscripten-memory64/LLGo 136047 B 0 B / +0.0% 91328 B 0 B / +0.0%
reflectcall/j32-emscripten/LLGo 1481123 B 0 B / +0.0% 104917 B 0 B / +0.0%
reflectcall/j32-goos-js/Go 2191221 B 0 B / +0.0% 0 B 0 B / 0.0%
reflectcall/j32-goos-js/LLGo 1530979 B 0 B / +0.0% 89743 B 0 B / +0.0%
reflectcall/j64-emscripten-memory64/LLGo 1373248 B 0 B / +0.0% 109492 B 0 B / +0.0%
reflectcall/w32-goos-wasip1/Go 2205707 B 0 B / +0.0% 0 B 0 B / 0.0%
reflectcall/w32-goos-wasip1/LLGo 1283895 B 0 B / +0.0% 0 B 0 B / 0.0%
reflectcall/w32-wasi/LLGo 1281420 B 0 B / +0.0% 0 B 0 B / 0.0%
w32-goos-wasip1/Go 1909947 B 0 B / +0.0% 0 B 0 B / 0.0%
w32-goos-wasip1/LLGo 153943 B 0 B / +0.0% 0 B 0 B / 0.0%
w32-wasi/LLGo 154626 B 0 B / +0.0% 0 B 0 B / 0.0%
LLGo WebAssembly build measurements
Example and profile Build vs base
j32-emscripten 7.174 s +226.2 ms / +3.3% (worse)
j32-goos-js 6.891 s -98.22 ms / -1.4% (better)
j64-emscripten-memory64 6.115 s -54.74 ms / -0.9% (better)
reflectcall/w32-wasi 23.597 s +138.9 ms / +0.6% (worse)
w32-goos-wasip1 4.267 s -148.6 ms / -3.4% (better)
w32-wasi 4.050 s -98.83 ms / -2.4% (better)

Compared with 1866f43f5b83 measured in the same runner job.

@github-actions

github-actions Bot commented Oct 6, 2026 •

Copy link
Copy Markdown

LLGo baseline benchmarks

6cb2cbe98e0c | workflow run | long-term charts

Program measurements

Platform Workload File size vs base Text size vs base Build vs base Run vs base
Linux cprintf 8568 B 0 B / +0.0% 387 B 0 B / +0.0% 955.164 ms +11.06 ms / +1.2% (worse) 1.235 ms -19.71 us / -1.6% (better)
Linux cprintf-lto 8408 B 0 B / +0.0% 368 B 0 B / +0.0% 940.826 ms -128.2 ms / -12.0% (better) 1.246 ms -8.694 us / -0.7% (better)
Linux fmtprintf 4886280 B 0 B / +0.0% 504761 B 0 B / +0.0% 7.338 s -238.8 ms / -3.2% (better) 3.089 ms -174.3 us / -5.3% (better)
Linux fmtprintf-lto 3616568 B 0 B / +0.0% 443075 B 0 B / +0.0% 16.419 s -150.6 ms / -0.9% (better) 2.936 ms -18.7 us / -0.6% (better)
Linux println 675040 B 0 B / +0.0% 16855 B 0 B / +0.0% 985.320 ms -46.9 ms / -4.5% (better) 1.641 ms -44.71 us / -2.7% (better)
Linux println-lto 190040 B 0 B / +0.0% 14273 B 0 B / +0.0% 1.266 s -83.9 ms / -6.2% (better) 1.580 ms -178.1 us / -10.1% (better)
macOS cprintf 50736 B 0 B / +0.0% 4409 B 0 B / +0.0% 840.200 ms -38.75 ms / -4.4% (better) 2.831 ms -326.2 us / -10.3% (better)
macOS cprintf-lto 50496 B 0 B / +0.0% 161 B 0 B / +0.0% 874.468 ms -202.5 ms / -18.8% (better) 3.106 ms -1.813 ms / -36.9% (better)
macOS fmtprintf 1774368 B 0 B / +0.0% 880529 B 0 B / +0.0% 4.293 s -318.8 ms / -6.9% (better) 8.843 ms -792.9 us / -8.2% (better)
macOS fmtprintf-lto 1361232 B 0 B / +0.0% 764957 B 0 B / +0.0% 12.892 s +2.699 s / +26.5% (worse) 9.239 ms +5.185 ms / +127.9% (worse)
macOS println 99344 B 0 B / +0.0% 24216 B 0 B / +0.0% 937.662 ms +3.466 ms / +0.4% (worse) 3.744 ms +130 us / +3.6% (worse)
macOS println-lto 83664 B 0 B / +0.0% 21457 B 0 B / +0.0% 1.112 s +13.17 ms / +1.2% (worse) 5.993 ms +2.592 ms / +76.2% (worse)
Windows MinGW cprintf 651264 B 0 B / +0.0% 4550 B 0 B / +0.0% 1.725 s -76.64 ms / -4.3% (better) 3.600 ms -259.2 us / -6.7% (better)
Windows MinGW cprintf-lto 44032 B 0 B / +0.0% 4502 B 0 B / +0.0% 1.853 s +58.21 ms / +3.2% (worse) 3.376 ms -928.7 us / -21.6% (better)
Windows MinGW fmtprintf 5440512 B 0 B / +0.0% 604470 B 0 B / +0.0% 7.406 s -25.92 ms / -0.3% (better) 8.573 ms +497.6 us / +6.2% (worse)
Windows MinGW fmtprintf-lto 4144128 B 0 B / +0.0% 551590 B 0 B / +0.0% 14.970 s -66.38 ms / -0.4% (better) 7.994 ms -93.6 us / -1.2% (better)
Windows MinGW println 706048 B 0 B / +0.0% 25190 B 0 B / +0.0% 1.720 s -27.38 ms / -1.6% (better) 6.571 ms -299.7 us / -4.4% (better)
Windows MinGW println-lto 208896 B 0 B / +0.0% 22070 B 0 B / +0.0% 1.992 s -57.15 ms / -2.8% (better) 6.365 ms -817.3 us / -11.4% (better)
Windows MinGW 386 cprintf 601600 B 0 B / +0.0% 5352 B 0 B / +0.0% 1.770 s +37.03 ms / +2.1% (worse) 5.134 ms -10.5 us / -0.2% (better)
Windows MinGW 386 cprintf-lto 103424 B 0 B / +0.0% 5112 B 0 B / +0.0% 1.801 s +94.63 ms / +5.5% (worse) 8.455 ms +3.318 ms / +64.6% (worse)
Windows MinGW 386 fmtprintf 4745728 B 0 B / +0.0% 472472 B 0 B / +0.0% 7.806 s +404.2 ms / +5.5% (worse) 11.325 ms +192.1 us / +1.7% (worse)
Windows MinGW 386 fmtprintf-lto 4148736 B 0 B / +0.0% 451132 B 0 B / +0.0% 15.023 s +434.5 ms / +3.0% (worse) 11.283 ms -299.7 us / -2.6% (better)
Windows MinGW 386 println 654336 B 0 B / +0.0% 21500 B 0 B / +0.0% 1.805 s +73.95 ms / +4.3% (worse) 9.676 ms +879.6 us / +10.0% (worse)
Windows MinGW 386 println-lto 258560 B 0 B / +0.0% 19328 B 0 B / +0.0% 2.062 s +76.02 ms / +3.8% (worse) 9.694 ms +1.211 ms / +14.3% (worse)
Windows MinGW ARM64 cprintf 662016 B 0 B / +0.0% 4436 B 0 B / +0.0% 1.811 s -42.29 ms / -2.3% (better) 5.797 ms -180.7 us / -3.0% (better)
Windows MinGW ARM64 cprintf-lto 43520 B 0 B / +0.0% 4388 B 0 B / +0.0% 1.838 s -89.22 ms / -4.6% (better) 6.026 ms -648 us / -9.7% (better)
Windows MinGW ARM64 fmtprintf 5348352 B 0 B / +0.0% 512664 B 0 B / +0.0% 6.838 s -38.27 ms / -0.6% (better) 12.143 ms -340.2 us / -2.7% (better)
Windows MinGW ARM64 fmtprintf-lto 4311552 B 0 B / +0.0% 478852 B 0 B / +0.0% 13.330 s -200 ms / -1.5% (better) 12.561 ms +77.5 us / +0.6% (worse)
Windows MinGW ARM64 println 713728 B 0 B / +0.0% 23912 B 0 B / +0.0% 1.833 s -28.09 ms / -1.5% (better) 10.397 ms +48.2 us / +0.5% (worse)
Windows MinGW ARM64 println-lto 215552 B 0 B / +0.0% 21272 B 0 B / +0.0% 2.068 s -60.13 ms / -2.8% (better) 10.164 ms -352.6 us / -3.4% (better)
Windows MSVC cprintf 895488 B 0 B / +0.0% 66019 B 0 B / +0.0% 1.521 s -119.3 ms / -7.3% (better) 3.349 ms -1.201 ms / -26.4% (better)
Windows MSVC cprintf-lto 291328 B 0 B / +0.0% 65971 B 0 B / +0.0% 1.591 s +37.57 ms / +2.4% (worse) 3.144 ms -254.3 us / -7.5% (better)
Windows MSVC fmtprintf 5742080 B 0 B / +0.0% 700214 B 0 B / +0.0% 7.149 s +200.2 ms / +2.9% (worse) 9.864 ms +858.6 us / +9.5% (worse)
Windows MSVC fmtprintf-lto 4469248 B 0 B / +0.0% 651350 B 0 B / +0.0% 13.650 s +93.97 ms / +0.7% (worse) 8.920 ms +254.5 us / +2.9% (worse)
Windows MSVC println 1018880 B 0 B / +0.0% 121062 B 0 B / +0.0% 1.555 s +26.38 ms / +1.7% (worse) 6.731 ms -160.9 us / -2.3% (better)
Windows MSVC println-lto 529920 B 0 B / +0.0% 118598 B 0 B / +0.0% 1.791 s +38.92 ms / +2.2% (worse) 6.722 ms -147.6 us / -2.1% (better)
Windows MSVC 386 cprintf 514048 B 0 B / +0.0% 3947 B 0 B / +0.0% 1.511 s -284.5 ms / -15.8% (better) 5.138 ms -629 us / -10.9% (better)
Windows MSVC 386 cprintf-lto 44032 B 0 B / +0.0% 3901 B 0 B / +0.0% 1.531 s -60.47 ms / -3.8% (better) 5.590 ms -306.6 us / -5.2% (better)
Windows MSVC 386 fmtprintf 4480000 B 0 B / +0.0% 455900 B 0 B / +0.0% 7.054 s -133.3 ms / -1.9% (better) 11.317 ms +283.6 us / +2.6% (worse)
Windows MSVC 386 fmtprintf-lto 3899904 B 0 B / +0.0% 426667 B 0 B / +0.0% 13.305 s -153.6 ms / -1.1% (better) 11.589 ms +218.7 us / +1.9% (worse)
Windows MSVC 386 println 568320 B 0 B / +0.0% 20356 B 0 B / +0.0% 1.519 s -41.54 ms / -2.7% (better) 9.524 ms -406.5 us / -4.1% (better)
Windows MSVC 386 println-lto 199168 B 0 B / +0.0% 18517 B 0 B / +0.0% 1.978 s +188.9 ms / +10.6% (worse) 9.356 ms +251.8 us / +2.8% (worse)
Windows MSVC ARM64 cprintf 662528 B 0 B / +0.0% 4240 B 0 B / +0.0% 1.667 s +17.22 ms / +1.0% (worse) 6.938 ms -39.5 us / -0.6% (better)
Windows MSVC ARM64 cprintf-lto 48640 B 0 B / +0.0% 4140 B 0 B / +0.0% 1.689 s +41.59 ms / +2.5% (worse) 7.244 ms +593.8 us / +8.9% (worse)
Windows MSVC ARM64 fmtprintf 5345792 B 0 B / +0.0% 512612 B 0 B / +0.0% 6.785 s +23.18 ms / +0.3% (worse) 13.647 ms -335.5 us / -2.4% (better)
Windows MSVC ARM64 fmtprintf-lto 4319232 B 0 B / +0.0% 479536 B 0 B / +0.0% 12.962 s -211.3 ms / -1.6% (better) 14.925 ms +43.9 us / +0.3% (worse)
Windows MSVC ARM64 println 716288 B 0 B / +0.0% 23956 B 0 B / +0.0% 1.624 s -29.26 ms / -1.8% (better) 12.189 ms -765.2 us / -5.9% (better)
Windows MSVC ARM64 println-lto 220672 B 0 B / +0.0% 21444 B 0 B / +0.0% 1.828 s -19.35 ms / -1.0% (better) 11.704 ms +543.2 us / +4.9% (worse)
Core language and compiler benchmarks
Platform Benchmark ns/op vs base
Linux BenchmarkLookupPCRandom 14.560 ns/op -0.08 ns/op / -0.5% (better)
Linux BenchmarkMergeCompilerFlags 192.800 ns/op +0.2 ns/op / +0.1% (worse)
Linux BenchmarkMergeLinkerFlags 127.800 ns/op +2.1 ns/op / +1.7% (worse)
Linux BenchmarkChannelBuffered 54.610 ns/op -0.15 ns/op / -0.3% (better)
Linux BenchmarkChannelHandoff 13100 ns/op +31 ns/op / +0.2% (worse)
Linux BenchmarkDefer 48.020 ns/op -1.06 ns/op / -2.2% (better)
Linux BenchmarkDirectCall 1.519 ns/op +0.002 ns/op / +0.1% (worse)
Linux BenchmarkGlobalRead 1.569 ns/op +0.006 ns/op / +0.4% (worse)
Linux BenchmarkGlobalWrite 7.765 ns/op -0.009 ns/op / -0.1% (better)
Linux BenchmarkGoroutine 23831 ns/op -142 ns/op / -0.6% (better)
Linux BenchmarkInterfaceCall 6.432 ns/op +0.212 ns/op / +3.4% (worse)
Linux BenchmarkRuntimeGetG 2.839 ns/op -0.025 ns/op / -0.9% (better)
macOS BenchmarkLookupPCRandom 13.430 ns/op +1.08 ns/op / +8.7% (worse)
macOS BenchmarkMergeCompilerFlags 126.100 ns/op +16.7 ns/op / +15.3% (worse)
macOS BenchmarkMergeLinkerFlags 70.290 ns/op +1.2 ns/op / +1.7% (worse)
macOS BenchmarkChannelBuffered 34.650 ns/op +8.94 ns/op / +34.8% (worse)
macOS BenchmarkChannelHandoff 8564 ns/op -967 ns/op / -10.1% (better)
macOS BenchmarkDefer 53.480 ns/op +19.98 ns/op / +59.6% (worse)
macOS BenchmarkDirectCall 1.284 ns/op +0.121 ns/op / +10.4% (worse)
macOS BenchmarkGlobalRead 1.232 ns/op +0.108 ns/op / +9.6% (worse)
macOS BenchmarkGlobalWrite 1.280 ns/op +0.152 ns/op / +13.5% (worse)
macOS BenchmarkGoroutine 73875 ns/op +35361 ns/op / +91.8% (worse)
macOS BenchmarkInterfaceCall 4.784 ns/op +0.714 ns/op / +17.5% (worse)
macOS BenchmarkRuntimeGetG 2.481 ns/op +0.283 ns/op / +12.9% (worse)
Windows MinGW BenchmarkLookupPCRandom 13.270 ns/op +0.02 ns/op / +0.2% (worse)
Windows MinGW BenchmarkMergeCompilerFlags 605.100 ns/op -16.6 ns/op / -2.7% (better)
Windows MinGW BenchmarkMergeLinkerFlags 533.200 ns/op -14.8 ns/op / -2.7% (better)
Windows MinGW BenchmarkChannelBuffered 30.740 ns/op -0.43 ns/op / -1.4% (better)
Windows MinGW BenchmarkChannelHandoff 873.800 ns/op -22.2 ns/op / -2.5% (better)
Windows MinGW BenchmarkDefer 57.210 ns/op +0.54 ns/op / +1.0% (worse)
Windows MinGW BenchmarkDirectCall 1.549 ns/op +0.001 ns/op / +0.1% (worse)
Windows MinGW BenchmarkGlobalRead 1.546 ns/op -0.002 ns/op / -0.1% (better)
Windows MinGW BenchmarkGlobalWrite 2.472 ns/op -0.001 ns/op / -0.04044% (better)
Windows MinGW BenchmarkGoroutine 91938 ns/op +286 ns/op / +0.3% (worse)
Windows MinGW BenchmarkInterfaceCall 7.754 ns/op -0.006 ns/op / -0.1% (better)
Windows MinGW BenchmarkRuntimeGetG 2.480 ns/op -0.001 ns/op / -0.04031% (better)
Windows MinGW 386 BenchmarkLookupPCRandom 26.640 ns/op +0.16 ns/op / +0.6% (worse)
Windows MinGW 386 BenchmarkMergeCompilerFlags 755.200 ns/op -2.3 ns/op / -0.3% (better)
Windows MinGW 386 BenchmarkMergeLinkerFlags 707.100 ns/op +2.6 ns/op / +0.4% (worse)
Windows MinGW 386 BenchmarkChannelBuffered 39.690 ns/op 0 ns/op / +0.0%
Windows MinGW 386 BenchmarkChannelHandoff 714.500 ns/op -222.1 ns/op / -23.7% (better)
Windows MinGW 386 BenchmarkDefer 44.030 ns/op -0.07 ns/op / -0.2% (better)
Windows MinGW 386 BenchmarkDirectCall 1.552 ns/op +0.004 ns/op / +0.3% (worse)
Windows MinGW 386 BenchmarkGlobalRead 1.870 ns/op +0.011 ns/op / +0.6% (worse)
Windows MinGW 386 BenchmarkGlobalWrite 7.784 ns/op +0.013 ns/op / +0.2% (worse)
Windows MinGW 386 BenchmarkGoroutine 108199 ns/op -481 ns/op / -0.4% (better)
Windows MinGW 386 BenchmarkInterfaceCall 8.141 ns/op +0.001 ns/op / +0.01229% (worse)
Windows MinGW 386 BenchmarkRuntimeGetG 2.167 ns/op -0.005 ns/op / -0.2% (better)
Windows MinGW ARM64 BenchmarkLookupPCRandom 12.050 ns/op -0.03 ns/op / -0.2% (better)
Windows MinGW ARM64 BenchmarkMergeCompilerFlags 569.300 ns/op -16.6 ns/op / -2.8% (better)
Windows MinGW ARM64 BenchmarkMergeLinkerFlags 530.800 ns/op -14.4 ns/op / -2.6% (better)
Windows MinGW ARM64 BenchmarkChannelBuffered 37.480 ns/op +0.02 ns/op / +0.1% (worse)
Windows MinGW ARM64 BenchmarkChannelHandoff 2102 ns/op +484 ns/op / +29.9% (worse)
Windows MinGW ARM64 BenchmarkDefer 53.200 ns/op -0.13 ns/op / -0.2% (better)
Windows MinGW ARM64 BenchmarkDirectCall 0.663 ns/op -0.0001 ns/op / -0.01508% (better)
Windows MinGW ARM64 BenchmarkGlobalRead 0.590 ns/op -0.0002 ns/op / -0.03391% (better)
Windows MinGW ARM64 BenchmarkGlobalWrite 0.737 ns/op -0.0002 ns/op / -0.02715% (better)
Windows MinGW ARM64 BenchmarkGoroutine 59204 ns/op -1839 ns/op / -3.0% (better)
Windows MinGW ARM64 BenchmarkInterfaceCall 4.149 ns/op -0.001 ns/op / -0.0241% (better)
Windows MinGW ARM64 BenchmarkRuntimeGetG 1.805 ns/op -0.001 ns/op / -0.1% (better)
Windows MSVC BenchmarkLookupPCRandom 12.950 ns/op -0.28 ns/op / -2.1% (better)
Windows MSVC BenchmarkMergeCompilerFlags 618.100 ns/op +0.5 ns/op / +0.1% (worse)
Windows MSVC BenchmarkMergeLinkerFlags 541.100 ns/op -6.1 ns/op / -1.1% (better)
Windows MSVC BenchmarkChannelBuffered 28.730 ns/op +0.02 ns/op / +0.1% (worse)
Windows MSVC BenchmarkChannelHandoff 1059 ns/op -24 ns/op / -2.2% (better)
Windows MSVC BenchmarkDefer 54.300 ns/op -0.2 ns/op / -0.4% (better)
Windows MSVC BenchmarkDirectCall 1.547 ns/op 0 ns/op / +0.0%
Windows MSVC BenchmarkGlobalRead 1.547 ns/op -0.001 ns/op / -0.1% (better)
Windows MSVC BenchmarkGlobalWrite 2.474 ns/op +0.007 ns/op / +0.3% (worse)
Windows MSVC BenchmarkGoroutine 90647 ns/op +1265 ns/op / +1.4% (worse)
Windows MSVC BenchmarkInterfaceCall 8.382 ns/op -0.002 ns/op / -0.02385% (better)
Windows MSVC BenchmarkRuntimeGetG 2.186 ns/op +0.014 ns/op / +0.6% (worse)
Windows MSVC 386 BenchmarkLookupPCRandom 26.520 ns/op +0.01 ns/op / +0.03772% (worse)
Windows MSVC 386 BenchmarkMergeCompilerFlags 784.900 ns/op +21 ns/op / +2.7% (worse)
Windows MSVC 386 BenchmarkMergeLinkerFlags 678.100 ns/op -18.4 ns/op / -2.6% (better)
Windows MSVC 386 BenchmarkChannelBuffered 39.270 ns/op +0.05 ns/op / +0.1% (worse)
Windows MSVC 386 BenchmarkChannelHandoff 755.400 ns/op -95.3 ns/op / -11.2% (better)
Windows MSVC 386 BenchmarkDefer 47.100 ns/op +1.38 ns/op / +3.0% (worse)
Windows MSVC 386 BenchmarkDirectCall 1.548 ns/op +0.001 ns/op / +0.1% (worse)
Windows MSVC 386 BenchmarkGlobalRead 1.553 ns/op +0.005 ns/op / +0.3% (worse)
Windows MSVC 386 BenchmarkGlobalWrite 7.792 ns/op +0.017 ns/op / +0.2% (worse)
Windows MSVC 386 BenchmarkGoroutine 112242 ns/op +125 ns/op / +0.1% (worse)
Windows MSVC 386 BenchmarkInterfaceCall 8.127 ns/op -0.013 ns/op / -0.2% (better)
Windows MSVC 386 BenchmarkRuntimeGetG 1.928 ns/op -0.001 ns/op / -0.1% (better)
Windows MSVC ARM64 BenchmarkLookupPCRandom 12.080 ns/op -0.06 ns/op / -0.5% (better)
Windows MSVC ARM64 BenchmarkMergeCompilerFlags 569.800 ns/op -1.7 ns/op / -0.3% (better)
Windows MSVC ARM64 BenchmarkMergeLinkerFlags 526.300 ns/op -4.8 ns/op / -0.9% (better)
Windows MSVC ARM64 BenchmarkChannelBuffered 38.630 ns/op -0.02 ns/op / -0.1% (better)
Windows MSVC ARM64 BenchmarkChannelHandoff 2559 ns/op +215 ns/op / +9.2% (worse)
Windows MSVC ARM64 BenchmarkDefer 64.850 ns/op +1.89 ns/op / +3.0% (worse)
Windows MSVC ARM64 BenchmarkDirectCall 0.590 ns/op -0.0004 ns/op / -0.1% (better)
Windows MSVC ARM64 BenchmarkGlobalRead 0.590 ns/op +0.0006 ns/op / +0.1% (worse)
Windows MSVC ARM64 BenchmarkGlobalWrite 3.832 ns/op -0.002 ns/op / -0.1% (better)
Windows MSVC ARM64 BenchmarkGoroutine 56798 ns/op +548 ns/op / +1.0% (worse)
Windows MSVC ARM64 BenchmarkInterfaceCall 4.152 ns/op +0.014 ns/op / +0.3% (worse)
Windows MSVC ARM64 BenchmarkRuntimeGetG 1.770 ns/op -0.039 ns/op / -2.2% (better)
Timer runtime benchmarks
Platform Operation and runtime ns/op vs base
Linux AfterFuncZeroDelivery/Go 913.500 ns/op +15.4 ns/op / +1.7% (worse)
Linux AfterFuncZeroDelivery/LLGo 40352 ns/op -3934 ns/op / -8.9% (better)
Linux CreateStop/Go 288.500 ns/op 0 ns/op / +0.0%
Linux CreateStop/LLGo 1644 ns/op -118 ns/op / -6.7% (better)
Linux RearmStopped/Go 116 ns/op 0 ns/op / +0.0%
Linux RearmStopped/LLGo 1386 ns/op +112 ns/op / +8.8% (worse)
Linux ResetActive/Go 68.600 ns/op -0.04 ns/op / -0.1% (better)
Linux ResetActive/LLGo 827.600 ns/op +97.4 ns/op / +13.3% (worse)
Linux ResetHeap1024/Go 67.110 ns/op +0.01 ns/op / +0.0149% (worse)
Linux ResetHeap1024/LLGo 178 ns/op -0.6 ns/op / -0.3% (better)
macOS AfterFuncZeroDelivery/Go 476.500 ns/op +49.3 ns/op / +11.5% (worse)
macOS AfterFuncZeroDelivery/LLGo 98762 ns/op +30842 ns/op / +45.4% (worse)
macOS CreateStop/Go 168.600 ns/op +32.5 ns/op / +23.9% (worse)
macOS CreateStop/LLGo 559.600 ns/op +134.5 ns/op / +31.6% (worse)
macOS RearmStopped/Go 64.520 ns/op +0.68 ns/op / +1.1% (worse)
macOS RearmStopped/LLGo 400.800 ns/op +71.2 ns/op / +21.6% (worse)
macOS ResetActive/Go 47.290 ns/op +4.8 ns/op / +11.3% (worse)
macOS ResetActive/LLGo 217.600 ns/op +54.8 ns/op / +33.7% (worse)
macOS ResetHeap1024/Go 48.220 ns/op +5.94 ns/op / +14.0% (worse)
macOS ResetHeap1024/LLGo 96.280 ns/op +9.57 ns/op / +11.0% (worse)
Windows MinGW AfterFuncZeroDelivery/Go 546.600 ns/op -24.4 ns/op / -4.3% (better)
Windows MinGW AfterFuncZeroDelivery/LLGo 185989 ns/op +324 ns/op / +0.2% (worse)
Windows MinGW CreateStop/Go 116.100 ns/op -0.2 ns/op / -0.2% (better)
Windows MinGW CreateStop/LLGo 417.400 ns/op -36.3 ns/op / -8.0% (better)
Windows MinGW RearmStopped/Go 31.140 ns/op -0.19 ns/op / -0.6% (better)
Windows MinGW RearmStopped/LLGo 272.500 ns/op -8.1 ns/op / -2.9% (better)
Windows MinGW ResetActive/Go 20.130 ns/op +0.03 ns/op / +0.1% (worse)
Windows MinGW ResetActive/LLGo 156 ns/op +0.8 ns/op / +0.5% (worse)
Windows MinGW ResetHeap1024/Go 20.360 ns/op -0.06 ns/op / -0.3% (better)
Windows MinGW ResetHeap1024/LLGo 126.300 ns/op -2.5 ns/op / -1.9% (better)
Windows MinGW 386 AfterFuncZeroDelivery/Go 964.200 ns/op +9.7 ns/op / +1.0% (worse)
Windows MinGW 386 AfterFuncZeroDelivery/LLGo 200692 ns/op +2162 ns/op / +1.1% (worse)
Windows MinGW 386 CreateStop/Go 194.900 ns/op +2.9 ns/op / +1.5% (worse)
Windows MinGW 386 CreateStop/LLGo 492.100 ns/op -4.6 ns/op / -0.9% (better)
Windows MinGW 386 RearmStopped/Go 63.650 ns/op +0.18 ns/op / +0.3% (worse)
Windows MinGW 386 RearmStopped/LLGo 352.800 ns/op +3.9 ns/op / +1.1% (worse)
Windows MinGW 386 ResetActive/Go 39.150 ns/op 0 ns/op / +0.0%
Windows MinGW 386 ResetActive/LLGo 996.600 ns/op -33.4 ns/op / -3.2% (better)
Windows MinGW 386 ResetHeap1024/Go 39.430 ns/op -0.05 ns/op / -0.1% (better)
Windows MinGW 386 ResetHeap1024/LLGo 190.900 ns/op +4.7 ns/op / +2.5% (worse)
Windows MinGW ARM64 AfterFuncZeroDelivery/Go 670.200 ns/op +12.3 ns/op / +1.9% (worse)
Windows MinGW ARM64 AfterFuncZeroDelivery/LLGo 140394 ns/op -2984 ns/op / -2.1% (better)
Windows MinGW ARM64 CreateStop/Go 200.500 ns/op +5.9 ns/op / +3.0% (worse)
Windows MinGW ARM64 CreateStop/LLGo 407.400 ns/op -18.2 ns/op / -4.3% (better)
Windows MinGW ARM64 RearmStopped/Go 70.590 ns/op -0.05 ns/op / -0.1% (better)
Windows MinGW ARM64 RearmStopped/LLGo 260.200 ns/op +1.8 ns/op / +0.7% (worse)
Windows MinGW ARM64 ResetActive/Go 30.880 ns/op -0.13 ns/op / -0.4% (better)
Windows MinGW ARM64 ResetActive/LLGo 139.400 ns/op +8.8 ns/op / +6.7% (worse)
Windows MinGW ARM64 ResetHeap1024/Go 31.070 ns/op -0.06 ns/op / -0.2% (better)
Windows MinGW ARM64 ResetHeap1024/LLGo 124 ns/op -0.5 ns/op / -0.4% (better)
Windows MSVC AfterFuncZeroDelivery/Go 572.700 ns/op +10.6 ns/op / +1.9% (worse)
Windows MSVC AfterFuncZeroDelivery/LLGo 174358 ns/op +762 ns/op / +0.4% (worse)
Windows MSVC CreateStop/Go 115.700 ns/op 0 ns/op / +0.0%
Windows MSVC CreateStop/LLGo 407.700 ns/op -9.2 ns/op / -2.2% (better)
Windows MSVC RearmStopped/Go 31.590 ns/op +0.29 ns/op / +0.9% (worse)
Windows MSVC RearmStopped/LLGo 253.500 ns/op -6.5 ns/op / -2.5% (better)
Windows MSVC ResetActive/Go 20.110 ns/op +0.1 ns/op / +0.5% (worse)
Windows MSVC ResetActive/LLGo 145 ns/op -8.6 ns/op / -5.6% (better)
Windows MSVC ResetHeap1024/Go 20.340 ns/op -0.11 ns/op / -0.5% (better)
Windows MSVC ResetHeap1024/LLGo 125.200 ns/op -3 ns/op / -2.3% (better)
Windows MSVC 386 AfterFuncZeroDelivery/Go 946.200 ns/op -4.4 ns/op / -0.5% (better)
Windows MSVC 386 AfterFuncZeroDelivery/LLGo 207629 ns/op -1556 ns/op / -0.7% (better)
Windows MSVC 386 CreateStop/Go 192 ns/op -1 ns/op / -0.5% (better)
Windows MSVC 386 CreateStop/LLGo 466.800 ns/op -23.7 ns/op / -4.8% (better)
Windows MSVC 386 RearmStopped/Go 63.430 ns/op -0.19 ns/op / -0.3% (better)
Windows MSVC 386 RearmStopped/LLGo 320.300 ns/op +4 ns/op / +1.3% (worse)
Windows MSVC 386 ResetActive/Go 39.070 ns/op -0.02 ns/op / -0.1% (better)
Windows MSVC 386 ResetActive/LLGo 997.800 ns/op -8.2 ns/op / -0.8% (better)
Windows MSVC 386 ResetHeap1024/Go 39.420 ns/op -0.03 ns/op / -0.1% (better)
Windows MSVC 386 ResetHeap1024/LLGo 173.500 ns/op -1.1 ns/op / -0.6% (better)
Windows MSVC ARM64 AfterFuncZeroDelivery/Go 674 ns/op +8.5 ns/op / +1.3% (worse)
Windows MSVC ARM64 AfterFuncZeroDelivery/LLGo 140734 ns/op -1118 ns/op / -0.8% (better)
Windows MSVC ARM64 CreateStop/Go 193.400 ns/op -4.5 ns/op / -2.3% (better)
Windows MSVC ARM64 CreateStop/LLGo 383.400 ns/op +2.9 ns/op / +0.8% (worse)
Windows MSVC ARM64 RearmStopped/Go 70.630 ns/op +0.04 ns/op / +0.1% (worse)
Windows MSVC ARM64 RearmStopped/LLGo 271.300 ns/op -0.7 ns/op / -0.3% (better)
Windows MSVC ARM64 ResetActive/Go 31.180 ns/op +0.12 ns/op / +0.4% (worse)
Windows MSVC ARM64 ResetActive/LLGo 134.200 ns/op +5.4 ns/op / +4.2% (worse)
Windows MSVC ARM64 ResetHeap1024/Go 31.170 ns/op +0.07 ns/op / +0.2% (worse)
Windows MSVC ARM64 ResetHeap1024/LLGo 137.400 ns/op +0.2 ns/op / +0.1% (worse)

Compared with 1866f43f5b83 measured in the same runner job.

@zhouguangyuan0718
zhouguangyuan0718 force-pushed the codex/js-simd-lto-20261006 branch from 4ba81bf to 6cb2cbe Compare October 7, 2026 00:01
@xushiwei
xushiwei merged commit d62e6aa into xgo-dev:main Oct 7, 2026
123 of 126 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants