Skip to content

wasm: unify browser exception handling on native SjLj - #2734

Draft
zhouguangyuan0718 wants to merge 5 commits into
xgo-dev:mainfrom
zhouguangyuan0718:codex/wasm-native-sjlj-20261005
Draft

zhouguangyuan0718 wants to merge 5 commits into
xgo-dev:mainfrom
zhouguangyuan0718:codex/wasm-native-sjlj-20261005

Conversation

@zhouguangyuan0718

@zhouguangyuan0718 zhouguangyuan0718 commented Oct 5, 2026 •

Copy link
Copy Markdown
Collaborator

Known compatibility regression; keep this PR draft: the previous JavaScript EH path supports the tested C++ catch -> Go callback -> Sleep -> resume sequence. Using the unmodified pre-change compiler ec9c2488b, both direct and indirect callbacks resumed and returned normally at O0 and O2 in Node and Chrome on Emscripten wasm32 (eight passing runs). The native Wasm EH/Asyncify path does not support suspension while the catch is active. The new negative-test assertions detect this loss of capability; they do not fix it. The default switch must be reconsidered against this regression before merge.

Browser targets currently lower Go panic/recover through Emscripten's JavaScript SjLj wrappers, while WASI uses native Wasm exceptions. JS wrappers cannot carry v128, requiring the SIMD call bridges and late-inlining restrictions in #2722. This PR selects native Wasm SjLj for GoJS, Emscripten wasm32, and Emscripten Memory64.

Compile and link with -sSUPPORT_LONGJMP=wasm -fwasm-exceptions. Both are needed: Emscripten bindings can reference C++ exception support even when the entry package contains only Go. The matching-LLVM Memory64 IR compiler uses native SjLj instead of -enable-emscripten-sjlj. Validate environment options and final package link arguments against incompatible SjLj settings and asyncify-ignore-unwind-from-catch.

Go's runtime retains its existing setjmp/longjmp implementation of defer/panic/recover. Browser builds retain Emscripten's legacy Wasm EH encoding for Asyncify; this is distinct from JavaScript SjLj. WASI retains its standard Wasm EH encoding. C++ exceptions must remain contained inside C++ wrappers and return a C ABI status to Go.

Qualification exposed three problems fixed here:

  • Forward Thin/Full LTO options to both compiler and linker drivers. Previously -lto changed the LLGo optimization pipeline without enabling browser-driver LTO.
  • Select SDK emar for Emscripten archives and MRI merging, preserving LLGO_AR overrides. Host LLVM 22 cannot index some C++ bitcode from the SDK's LLVM 24.
  • Reserve Wasm jump buffers in the function entry block. A defer introduced after control flow could previously allocate a dynamic shadow-stack buffer that a caught C++ exception discarded; subsequent suspension/GC overwrote its saved jump state and Go recovery failed. The expanded boundary fixture fails with the previous compiler and passes with this fix.

Asyncify has an explicit unsupported boundary: a callback must not suspend while a foreign Wasm catch is active. The Go/C++ fixture keeps its catch-status function behind a volatile function pointer because LLVM noinline alone does not preserve that boundary through Binaryen optimization. After returning from the wrapper, Go sleeps, collects, panics, and sleeps/collects again in the recovering defer.

The acceptance matrix enables existing Binaryen asyncify-asserts with Emscripten ASSERTIONS=0, and verifies that direct and indirect callbacks that deliberately suspend inside a C++ catch enter the callback, trap with unreachable, and never resume. These fixed negative callbacks are declared noexcept; nested cleanup/unwind shapes are not comprehensively qualified.

The extra assertions are only enabled for this focused acceptance matrix, not default application builds. The same Binaryen switch also rejects LLGo's intentional reflection trampoline suspension; TestReflectCallAndMethod passes with the default profile and traps with the extra assertions. There is no catch-only switch in the pinned Binaryen. This PR makes no new Binaryen or Emscripten patch, keeps the generic Asyncify/native-EH warning, and does not implement static analysis or automatic rejection of every unsupported catch/suspension path. Assertions detect unsupported behavior; they do not make it supported.

No SIMD implementation changes are included. This branch is based directly on main ec9c2488b, independently of #2722.

Validation on macOS arm64 with Go 1.27.0, LLGo LLVM 22.1.8, Emscripten 6.0.8, LLGo Binaryen llgo-v132.3, Node 24.19.0, Chrome, and Wasmer 7.5.0:

  • SSA, crosscompile, and clang package tests passed; focused build tests covered driver/profile selection, native EH, Memory64 IR compilation, archivers, and incompatible options.
  • Suspended Go panic/recover/Goexit passed all 12 combinations of three browser targets and O0/O2/actual Thin/Full LTO in both Node and Chrome. Acceptance checks native EH and absence of JS invoke_* imports.
  • Go/C++ boundary acceptance passed the same 12 combinations in both Node and Chrome, with 24 negative direct/indirect catch-suspension runs per engine. Actual LTO driver flags were verified.
  • Six standalone C++ EH comparison variants passed in Node and Chrome.
  • Complete single-worker acceptance passed: 27 checks/package runs, five expected failure paths, one timeout path, zero unexpected failures or skips.
  • Complete multi-worker acceptance passed in Node and Chrome, including runtime, GC, reflection, callbacks, and filesystem checks.
  • Target-profile acceptance passed, including WASI and raw GoJS.
  • Earlier, a separate integration checkout with simd: expand SIMD128 operations and runtime integration #2722 passed O0 SIMD boundary execution and the complete SIMD suite at O2 with actual Full LTO for all three browser targets while bypassing the entire SIMD bridge pass. That integration check preceded the latest fixed-frame/boundary additions; simd: expand SIMD128 operations and runtime integration #2722 itself was not modified.

Draft for review of the default-policy change. The pinned toolchain and explicit runtime matrix are required; this PR does not claim compatibility with arbitrary SDKs, arbitrary foreign catch/control-flow shapes, or foreign exceptions unwinding through suspended Go frames. CI remains in progress. The observed README link-check failure was a Codecov TLS HandshakeFailure, unrelated to these source changes; a green CI result would not resolve the compatibility regression above.

@github-actions

github-actions Bot commented Oct 5, 2026 •

Copy link
Copy Markdown

LLGo WebAssembly build benchmarks

e413103c3684 | workflow run | long-term charts

WebAssembly output sizes
Example, profile and compiler Wasm module vs base Generated JS glue vs base
cprintf/j32-emscripten/LLGo 149221 B -19 B / -0.01273% (better) 87966 B -191 B / -0.2% (better)
cprintf/j32-goos-js/LLGo 148492 B -19 B / -0.01279% (better) 72794 B -191 B / -0.3% (better)
cprintf/j64-emscripten-memory64/LLGo 136688 B -19 B / -0.0139% (better) 91086 B -242 B / -0.3% (better)
cprintf/w32-goos-wasip1/LLGo 153508 B 0 B / +0.0% 0 B 0 B / 0.0%
cprintf/w32-wasi/LLGo 154198 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/j32-emscripten/LLGo 3089986 B +40860 B / +1.3% (worse) 122374 B -7999 B / -6.1% (better)
fmtprintf/j32-goos-js/Go 2526852 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/j32-goos-js/LLGo 3184641 B +39489 B / +1.3% (worse) 93098 B -7675 B / -7.6% (better)
fmtprintf/j64-emscripten-memory64/LLGo 2831792 B +19950 B / +0.7% (worse) 130010 B -5843 B / -4.3% (better)
fmtprintf/w32-goos-wasip1/Go 2500019 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/w32-goos-wasip1/LLGo 2349407 B +71 B / +0.003022% (worse) 0 B 0 B / 0.0%
fmtprintf/w32-wasi/LLGo 2346039 B +63 B / +0.002685% (worse) 0 B 0 B / 0.0%
j32-emscripten/LLGo 148563 B -19 B / -0.01279% (better) 87966 B -191 B / -0.2% (better)
j32-goos-js/Go 1895533 B 0 B / +0.0% 0 B 0 B / 0.0%
j32-goos-js/LLGo 147983 B -19 B / -0.01284% (better) 72794 B -191 B / -0.3% (better)
j64-emscripten-memory64/LLGo 136028 B -19 B / -0.01397% (better) 91086 B -242 B / -0.3% (better)
reflectcall/j32-emscripten/LLGo 1470053 B -3206 B / -0.2% (better) 101479 B -3438 B / -3.3% (better)
reflectcall/j32-goos-js/Go 2191221 B 0 B / +0.0% 0 B 0 B / 0.0%
reflectcall/j32-goos-js/LLGo 1518536 B -4650 B / -0.3% (better) 86305 B -3438 B / -3.8% (better)
reflectcall/j64-emscripten-memory64/LLGo 1363090 B -2935 B / -0.2% (better) 105983 B -3509 B / -3.2% (better)
reflectcall/w32-goos-wasip1/Go 2205707 B 0 B / +0.0% 0 B 0 B / 0.0%
reflectcall/w32-goos-wasip1/LLGo 1277893 B +117 B / +0.009157% (worse) 0 B 0 B / 0.0%
reflectcall/w32-wasi/LLGo 1275471 B +109 B / +0.008547% (worse) 0 B 0 B / 0.0%
w32-goos-wasip1/Go 1909947 B 0 B / +0.0% 0 B 0 B / 0.0%
w32-goos-wasip1/LLGo 153155 B 0 B / +0.0% 0 B 0 B / 0.0%
w32-wasi/LLGo 153845 B 0 B / +0.0% 0 B 0 B / 0.0%
LLGo WebAssembly build measurements
Example and profile Build vs base
j32-emscripten 7.300 s +422.5 ms / +6.1% (worse)
j32-goos-js 7.015 s +257 ms / +3.8% (worse)
j64-emscripten-memory64 6.112 s -4.622 ms / -0.1% (better)
reflectcall/w32-wasi 23.119 s -217.5 ms / -0.9% (better)
w32-goos-wasip1 4.313 s -27.98 ms / -0.6% (better)
w32-wasi 4.024 s -59.52 ms / -1.5% (better)

Compared with ec9c2488bc37 measured in the same runner job.

@github-actions

github-actions Bot commented Oct 5, 2026 •

Copy link
Copy Markdown

LLGo baseline benchmarks

e413103c3684 | workflow run | long-term charts

Program measurements

Platform Workload File size vs base Text size vs base Build vs base Run vs base
Linux cprintf 8568 B 0 B / +0.0% 387 B 0 B / +0.0% 1.196 s -186.8 ms / -13.5% (better) 1.348 ms -56.85 us / -4.0% (better)
Linux cprintf-lto 8408 B 0 B / +0.0% 368 B 0 B / +0.0% 1.216 s +25.94 ms / +2.2% (worse) 1.348 ms -37.9 us / -2.7% (better)
Linux fmtprintf 4880944 B 0 B / +0.0% 501619 B 0 B / +0.0% 9.194 s -524.4 ms / -5.4% (better) 3.416 ms +72.5 us / +2.2% (worse)
Linux fmtprintf-lto 3600808 B 0 B / +0.0% 438538 B 0 B / +0.0% 18.707 s -872.9 ms / -4.5% (better) 3.175 ms -162.9 us / -4.9% (better)
Linux println 674952 B 0 B / +0.0% 16855 B 0 B / +0.0% 1.313 s +124.2 ms / +10.4% (worse) 1.732 ms -73.37 us / -4.1% (better)
Linux println-lto 190024 B 0 B / +0.0% 14273 B 0 B / +0.0% 1.628 s +178.5 ms / +12.3% (worse) 1.823 ms +7.587 us / +0.4% (worse)
macOS cprintf 50736 B 0 B / +0.0% 4409 B 0 B / +0.0% 872.737 ms -125.4 ms / -12.6% (better) 2.502 ms -360.3 us / -12.6% (better)
macOS cprintf-lto 50496 B 0 B / +0.0% 161 B 0 B / +0.0% 821.751 ms -274.7 ms / -25.1% (better) 3.018 ms -3.425 ms / -53.2% (better)
macOS fmtprintf 1773456 B 0 B / +0.0% 877780 B 0 B / +0.0% 4.968 s -104.3 ms / -2.1% (better) 8.796 ms -1.394 ms / -13.7% (better)
macOS fmtprintf-lto 1360928 B 0 B / +0.0% 762428 B 0 B / +0.0% 9.162 s -575.9 ms / -5.9% (better) 4.387 ms -103.1 us / -2.3% (better)
macOS println 99344 B 0 B / +0.0% 24216 B 0 B / +0.0% 846.289 ms -223.4 ms / -20.9% (better) 3.366 ms -171.9 us / -4.9% (better)
macOS println-lto 83664 B 0 B / +0.0% 21457 B 0 B / +0.0% 1.017 s -118.6 ms / -10.4% (better) 3.330 ms -14.46 ms / -81.3% (better)
Windows MinGW cprintf 651264 B 0 B / +0.0% 4550 B 0 B / +0.0% 1.730 s -132.6 us / -0.007666% (better) 3.540 ms -274.5 us / -7.2% (better)
Windows MinGW cprintf-lto 43520 B 0 B / +0.0% 4486 B 0 B / +0.0% 1.786 s +20.18 ms / +1.1% (worse) 3.503 ms -771 us / -18.0% (better)
Windows MinGW fmtprintf 5432832 B 0 B / +0.0% 600310 B 0 B / +0.0% 7.425 s +199.8 ms / +2.8% (worse) 9.122 ms +148 us / +1.6% (worse)
Windows MinGW fmtprintf-lto 4127232 B 0 B / +0.0% 546582 B 0 B / +0.0% 15.063 s -120.7 ms / -0.8% (better) 9.179 ms +1.313 ms / +16.7% (worse)
Windows MinGW println 706560 B 0 B / +0.0% 25190 B 0 B / +0.0% 1.762 s +27.31 ms / +1.6% (worse) 6.867 ms -74.5 us / -1.1% (better)
Windows MinGW println-lto 208896 B 0 B / +0.0% 22054 B 0 B / +0.0% 2.050 s +2.052 ms / +0.1% (worse) 7.161 ms +199.9 us / +2.9% (worse)
Windows MinGW 386 cprintf 601600 B 0 B / +0.0% 5326 B 0 B / +0.0% 1.748 s -30.51 ms / -1.7% (better) 5.196 ms -559.4 us / -9.7% (better)
Windows MinGW 386 cprintf-lto 103424 B 0 B / +0.0% 5094 B 0 B / +0.0% 1.765 s -93.24 ms / -5.0% (better) 5.163 ms -594.8 us / -10.3% (better)
Windows MinGW 386 fmtprintf 4745216 B 0 B / +0.0% 472478 B 0 B / +0.0% 7.667 s -16.73 ms / -0.2% (better) 11.148 ms +90.7 us / +0.8% (worse)
Windows MinGW 386 fmtprintf-lto 4148736 B 0 B / +0.0% 451114 B 0 B / +0.0% 14.946 s -258.7 ms / -1.7% (better) 10.621 ms -536.7 us / -4.8% (better)
Windows MinGW 386 println 653312 B 0 B / +0.0% 21490 B 0 B / +0.0% 1.793 s -33.23 ms / -1.8% (better) 9.386 ms -432.9 us / -4.4% (better)
Windows MinGW 386 println-lto 258560 B 0 B / +0.0% 19306 B 0 B / +0.0% 2.044 s -69.71 ms / -3.3% (better) 9.619 ms -213.5 us / -2.2% (better)
Windows MinGW ARM64 cprintf 661504 B 0 B / +0.0% 4408 B 0 B / +0.0% 1.887 s -5.449 ms / -0.3% (better) 6.276 ms -671.5 us / -9.7% (better)
Windows MinGW ARM64 cprintf-lto 43520 B 0 B / +0.0% 4340 B 0 B / +0.0% 1.921 s -18.09 ms / -0.9% (better) 6.342 ms -698.6 us / -9.9% (better)
Windows MinGW ARM64 fmtprintf 5343744 B 0 B / +0.0% 510876 B 0 B / +0.0% 7.137 s +14.87 ms / +0.2% (worse) 13.838 ms +813.8 us / +6.2% (worse)
Windows MinGW ARM64 fmtprintf-lto 4302336 B 0 B / +0.0% 477264 B 0 B / +0.0% 13.909 s +236.2 ms / +1.7% (worse) 13.501 ms +428.9 us / +3.3% (worse)
Windows MinGW ARM64 println 714240 B 0 B / +0.0% 23884 B 0 B / +0.0% 1.891 s -7.36 ms / -0.4% (better) 11.220 ms -175.6 us / -1.5% (better)
Windows MinGW ARM64 println-lto 215552 B 0 B / +0.0% 21232 B 0 B / +0.0% 2.160 s -12.35 ms / -0.6% (better) 11.255 ms +464.1 us / +4.3% (worse)
Windows MSVC cprintf 893952 B 0 B / +0.0% 65798 B 0 B / +0.0% 1.567 s -140.2 ms / -8.2% (better) 3.360 ms +18 us / +0.5% (worse)
Windows MSVC cprintf-lto 289792 B 0 B / +0.0% 65734 B 0 B / +0.0% 1.676 s +77.83 ms / +4.9% (worse) 4.962 ms +1.388 ms / +38.8% (worse)
Windows MSVC fmtprintf 5732864 B 0 B / +0.0% 695862 B 0 B / +0.0% 7.182 s +265.1 ms / +3.8% (worse) 9.231 ms +379.7 us / +4.3% (worse)
Windows MSVC fmtprintf-lto 4449792 B 0 B / +0.0% 646134 B 0 B / +0.0% 13.765 s +341.2 ms / +2.5% (worse) 9.293 ms +79.6 us / +0.9% (worse)
Windows MSVC println 1017344 B 0 B / +0.0% 120854 B 0 B / +0.0% 1.604 s -21.9 ms / -1.3% (better) 7.011 ms -542.8 us / -7.2% (better)
Windows MSVC println-lto 528384 B 0 B / +0.0% 118390 B 0 B / +0.0% 1.849 s +44.16 ms / +2.4% (worse) 7.556 ms +536.6 us / +7.6% (worse)
Windows MSVC 386 cprintf 513536 B 0 B / +0.0% 3931 B 0 B / +0.0% 1.134 s -193.5 ms / -14.6% (better) 4.365 ms -105.1 us / -2.4% (better)
Windows MSVC 386 cprintf-lto 44032 B 0 B / +0.0% 3853 B 0 B / +0.0% 1.139 s -78.21 ms / -6.4% (better) 3.966 ms -551.5 us / -12.2% (better)
Windows MSVC 386 fmtprintf 4479488 B 0 B / +0.0% 455868 B 0 B / +0.0% 5.659 s +66.46 ms / +1.2% (worse) 8.585 ms -304.4 us / -3.4% (better)
Windows MSVC 386 fmtprintf-lto 3899392 B 0 B / +0.0% 426651 B 0 B / +0.0% 10.622 s -302.7 ms / -2.8% (better) 8.455 ms -767.8 us / -8.3% (better)
Windows MSVC 386 println 567296 B 0 B / +0.0% 20340 B 0 B / +0.0% 1.149 s -55.27 ms / -4.6% (better) 7.308 ms -23.1 us / -0.3% (better)
Windows MSVC 386 println-lto 199168 B 0 B / +0.0% 18501 B 0 B / +0.0% 1.335 s -97.6 ms / -6.8% (better) 7.282 ms -442.7 us / -5.7% (better)
Windows MSVC ARM64 cprintf 662528 B 0 B / +0.0% 4192 B 0 B / +0.0% 1.612 s +948.3 us / +0.1% (worse) 6.917 ms +47.4 us / +0.7% (worse)
Windows MSVC ARM64 cprintf-lto 47616 B 0 B / +0.0% 4084 B 0 B / +0.0% 1.653 s +24.42 ms / +1.5% (worse) 6.903 ms +106 us / +1.6% (worse)
Windows MSVC ARM64 fmtprintf 5339648 B 0 B / +0.0% 510808 B 0 B / +0.0% 6.804 s +102.6 ms / +1.5% (worse) 14.873 ms +774.4 us / +5.5% (worse)
Windows MSVC ARM64 fmtprintf-lto 4309504 B 0 B / +0.0% 477924 B 0 B / +0.0% 13.230 s +136.4 ms / +1.0% (worse) 14.548 ms -562 us / -3.7% (better)
Windows MSVC ARM64 println 715264 B 0 B / +0.0% 23908 B 0 B / +0.0% 1.621 s -3.262 ms / -0.2% (better) 11.705 ms -515.6 us / -4.2% (better)
Windows MSVC ARM64 println-lto 220672 B 0 B / +0.0% 21380 B 0 B / +0.0% 1.889 s +57.99 ms / +3.2% (worse) 13.213 ms +1.623 ms / +14.0% (worse)
Core language and compiler benchmarks
Platform Benchmark ns/op vs base
Linux BenchmarkLookupPCRandom 14.770 ns/op +0.04 ns/op / +0.3% (worse)
Linux BenchmarkMergeCompilerFlags 213 ns/op +7.7 ns/op / +3.8% (worse)
Linux BenchmarkMergeLinkerFlags 145.800 ns/op +11.8 ns/op / +8.8% (worse)
Linux BenchmarkChannelBuffered 55.170 ns/op +0.16 ns/op / +0.3% (worse)
Linux BenchmarkChannelHandoff 13433 ns/op -77 ns/op / -0.6% (better)
Linux BenchmarkDefer 48.100 ns/op -1.61 ns/op / -3.2% (better)
Linux BenchmarkDirectCall 1.169 ns/op +0.003 ns/op / +0.3% (worse)
Linux BenchmarkGlobalRead 1.164 ns/op -0.011 ns/op / -0.9% (better)
Linux BenchmarkGlobalWrite 7.785 ns/op +0.014 ns/op / +0.2% (worse)
Linux BenchmarkGoroutine 25863 ns/op -4830 ns/op / -15.7% (better)
Linux BenchmarkInterfaceCall 5.840 ns/op +0.014 ns/op / +0.2% (worse)
Linux BenchmarkRuntimeGetG 2.970 ns/op -0.036 ns/op / -1.2% (better)
macOS BenchmarkLookupPCRandom 12.620 ns/op +0.34 ns/op / +2.8% (worse)
macOS BenchmarkMergeCompilerFlags 89.280 ns/op -35.22 ns/op / -28.3% (better)
macOS BenchmarkMergeLinkerFlags 63.690 ns/op -3.54 ns/op / -5.3% (better)
macOS BenchmarkChannelBuffered 27.680 ns/op +0.96 ns/op / +3.6% (worse)
macOS BenchmarkChannelHandoff 6255 ns/op -3157 ns/op / -33.5% (better)
macOS BenchmarkDefer 33.660 ns/op -4.15 ns/op / -11.0% (better)
macOS BenchmarkDirectCall 1.148 ns/op +0.013 ns/op / +1.1% (worse)
macOS BenchmarkGlobalRead 1.033 ns/op -0.019 ns/op / -1.8% (better)
macOS BenchmarkGlobalWrite 1.031 ns/op +0.008 ns/op / +0.8% (worse)
macOS BenchmarkGoroutine 39710 ns/op +6200 ns/op / +18.5% (worse)
macOS BenchmarkInterfaceCall 3.971 ns/op -0.698 ns/op / -14.9% (better)
macOS BenchmarkRuntimeGetG 2.105 ns/op -0.117 ns/op / -5.3% (better)
Windows MinGW BenchmarkLookupPCRandom 12.320 ns/op -0.23 ns/op / -1.8% (better)
Windows MinGW BenchmarkMergeCompilerFlags 537.700 ns/op -4.3 ns/op / -0.8% (better)
Windows MinGW BenchmarkMergeLinkerFlags 472 ns/op +0.9 ns/op / +0.2% (worse)
Windows MinGW BenchmarkChannelBuffered 30.460 ns/op -0.03 ns/op / -0.1% (better)
Windows MinGW BenchmarkChannelHandoff 1326 ns/op -8 ns/op / -0.6% (better)
Windows MinGW BenchmarkDefer 55.570 ns/op -0.22 ns/op / -0.4% (better)
Windows MinGW BenchmarkDirectCall 1.748 ns/op +0.002 ns/op / +0.1% (worse)
Windows MinGW BenchmarkGlobalRead 1.749 ns/op +0.002 ns/op / +0.1% (worse)
Windows MinGW BenchmarkGlobalWrite 2.790 ns/op +0.002 ns/op / +0.1% (worse)
Windows MinGW BenchmarkGoroutine 77724 ns/op +82 ns/op / +0.1% (worse)
Windows MinGW BenchmarkInterfaceCall 8.772 ns/op +0.014 ns/op / +0.2% (worse)
Windows MinGW BenchmarkRuntimeGetG 2.449 ns/op +0.003 ns/op / +0.1% (worse)
Windows MinGW 386 BenchmarkLookupPCRandom 26.530 ns/op 0 ns/op / +0.0%
Windows MinGW 386 BenchmarkMergeCompilerFlags 727.800 ns/op -53.1 ns/op / -6.8% (better)
Windows MinGW 386 BenchmarkMergeLinkerFlags 704.200 ns/op -9.9 ns/op / -1.4% (better)
Windows MinGW 386 BenchmarkChannelBuffered 41.200 ns/op -0.23 ns/op / -0.6% (better)
Windows MinGW 386 BenchmarkChannelHandoff 869.400 ns/op +34.4 ns/op / +4.1% (worse)
Windows MinGW 386 BenchmarkDefer 41.810 ns/op -2.42 ns/op / -5.5% (better)
Windows MinGW 386 BenchmarkDirectCall 1.549 ns/op +0.001 ns/op / +0.1% (worse)
Windows MinGW 386 BenchmarkGlobalRead 1.550 ns/op -0.004 ns/op / -0.3% (better)
Windows MinGW 386 BenchmarkGlobalWrite 7.799 ns/op +0.015 ns/op / +0.2% (worse)
Windows MinGW 386 BenchmarkGoroutine 106656 ns/op -2075 ns/op / -1.9% (better)
Windows MinGW 386 BenchmarkInterfaceCall 8.371 ns/op -0.035 ns/op / -0.4% (better)
Windows MinGW 386 BenchmarkRuntimeGetG 1.924 ns/op -0.005 ns/op / -0.3% (better)
Windows MinGW ARM64 BenchmarkLookupPCRandom 12.170 ns/op +0.07 ns/op / +0.6% (worse)
Windows MinGW ARM64 BenchmarkMergeCompilerFlags 572.600 ns/op -4.9 ns/op / -0.8% (better)
Windows MinGW ARM64 BenchmarkMergeLinkerFlags 540 ns/op -39.9 ns/op / -6.9% (better)
Windows MinGW ARM64 BenchmarkChannelBuffered 38.940 ns/op +0.02 ns/op / +0.1% (worse)
Windows MinGW ARM64 BenchmarkChannelHandoff 2458 ns/op +138 ns/op / +5.9% (worse)
Windows MinGW ARM64 BenchmarkDefer 56.530 ns/op +1.98 ns/op / +3.6% (worse)
Windows MinGW ARM64 BenchmarkDirectCall 0.663 ns/op -0.0002 ns/op / -0.03014% (better)
Windows MinGW ARM64 BenchmarkGlobalRead 0.663 ns/op -0.0013 ns/op / -0.2% (better)
Windows MinGW ARM64 BenchmarkGlobalWrite 0.737 ns/op +0.0005 ns/op / +0.1% (worse)
Windows MinGW ARM64 BenchmarkGoroutine 69541 ns/op +5507 ns/op / +8.6% (worse)
Windows MinGW ARM64 BenchmarkInterfaceCall 4.144 ns/op +0.003 ns/op / +0.1% (worse)
Windows MinGW ARM64 BenchmarkRuntimeGetG 1.773 ns/op -0.03 ns/op / -1.7% (better)
Windows MSVC BenchmarkLookupPCRandom 13.190 ns/op +0.39 ns/op / +3.0% (worse)
Windows MSVC BenchmarkMergeCompilerFlags 594.400 ns/op -42.5 ns/op / -6.7% (better)
Windows MSVC BenchmarkMergeLinkerFlags 541 ns/op -14.2 ns/op / -2.6% (better)
Windows MSVC BenchmarkChannelBuffered 29.730 ns/op +0.04 ns/op / +0.1% (worse)
Windows MSVC BenchmarkChannelHandoff 1110 ns/op +35 ns/op / +3.3% (worse)
Windows MSVC BenchmarkDefer 53.870 ns/op -0.09 ns/op / -0.2% (better)
Windows MSVC BenchmarkDirectCall 1.550 ns/op +0.002 ns/op / +0.1% (worse)
Windows MSVC BenchmarkGlobalRead 1.547 ns/op -0.001 ns/op / -0.1% (better)
Windows MSVC BenchmarkGlobalWrite 2.467 ns/op -0.003 ns/op / -0.1% (better)
Windows MSVC BenchmarkGoroutine 90206 ns/op -420 ns/op / -0.5% (better)
Windows MSVC BenchmarkInterfaceCall 9.006 ns/op -0.025 ns/op / -0.3% (better)
Windows MSVC BenchmarkRuntimeGetG 2.485 ns/op +0.003 ns/op / +0.1% (worse)
Windows MSVC 386 BenchmarkLookupPCRandom 21.560 ns/op +0.02 ns/op / +0.1% (worse)
Windows MSVC 386 BenchmarkMergeCompilerFlags 545 ns/op -11.7 ns/op / -2.1% (better)
Windows MSVC 386 BenchmarkMergeLinkerFlags 526.300 ns/op +5.7 ns/op / +1.1% (worse)
Windows MSVC 386 BenchmarkChannelBuffered 33.930 ns/op 0 ns/op / +0.0%
Windows MSVC 386 BenchmarkChannelHandoff 713.800 ns/op +18.5 ns/op / +2.7% (worse)
Windows MSVC 386 BenchmarkDefer 40.530 ns/op +2.31 ns/op / +6.0% (worse)
Windows MSVC 386 BenchmarkDirectCall 1.357 ns/op -0.001 ns/op / -0.1% (better)
Windows MSVC 386 BenchmarkGlobalRead 1.358 ns/op -0.002 ns/op / -0.1% (better)
Windows MSVC 386 BenchmarkGlobalWrite 6.983 ns/op +0.002 ns/op / +0.02865% (worse)
Windows MSVC 386 BenchmarkGoroutine 73049 ns/op +1860 ns/op / +2.6% (worse)
Windows MSVC 386 BenchmarkInterfaceCall 7.335 ns/op -0.007 ns/op / -0.1% (better)
Windows MSVC 386 BenchmarkRuntimeGetG 1.631 ns/op -0.005 ns/op / -0.3% (better)
Windows MSVC ARM64 BenchmarkLookupPCRandom 12.010 ns/op 0 ns/op / +0.0%
Windows MSVC ARM64 BenchmarkMergeCompilerFlags 563.600 ns/op -28.5 ns/op / -4.8% (better)
Windows MSVC ARM64 BenchmarkMergeLinkerFlags 528.100 ns/op -21.9 ns/op / -4.0% (better)
Windows MSVC ARM64 BenchmarkChannelBuffered 39.490 ns/op +1.92 ns/op / +5.1% (worse)
Windows MSVC ARM64 BenchmarkChannelHandoff 2972 ns/op -741 ns/op / -20.0% (better)
Windows MSVC ARM64 BenchmarkDefer 62.870 ns/op +2.43 ns/op / +4.0% (worse)
Windows MSVC ARM64 BenchmarkDirectCall 0.663 ns/op 0 ns/op / +0.0%
Windows MSVC ARM64 BenchmarkGlobalRead 0.663 ns/op +0.0003 ns/op / +0.04524% (worse)
Windows MSVC ARM64 BenchmarkGlobalWrite 3.798 ns/op +0.003 ns/op / +0.1% (worse)
Windows MSVC ARM64 BenchmarkGoroutine 58026 ns/op +118 ns/op / +0.2% (worse)
Windows MSVC ARM64 BenchmarkInterfaceCall 4.140 ns/op -0.011 ns/op / -0.3% (better)
Windows MSVC ARM64 BenchmarkRuntimeGetG 1.809 ns/op +0.04 ns/op / +2.3% (worse)
Timer runtime benchmarks
Platform Operation and runtime ns/op vs base
Linux AfterFuncZeroDelivery/Go 912.300 ns/op -41.7 ns/op / -4.4% (better)
Linux AfterFuncZeroDelivery/LLGo 47546 ns/op +8783 ns/op / +22.7% (worse)
Linux CreateStop/Go 304.700 ns/op +2.1 ns/op / +0.7% (worse)
Linux CreateStop/LLGo 1486 ns/op -518 ns/op / -25.8% (better)
Linux RearmStopped/Go 115.100 ns/op -0.9 ns/op / -0.8% (better)
Linux RearmStopped/LLGo 1247 ns/op -193 ns/op / -13.4% (better)
Linux ResetActive/Go 67.870 ns/op -0.71 ns/op / -1.0% (better)
Linux ResetActive/LLGo 749 ns/op -60 ns/op / -7.4% (better)
Linux ResetHeap1024/Go 67.240 ns/op +0.13 ns/op / +0.2% (worse)
Linux ResetHeap1024/LLGo 184.500 ns/op +7.7 ns/op / +4.4% (worse)
macOS AfterFuncZeroDelivery/Go 365.900 ns/op -70 ns/op / -16.1% (better)
macOS AfterFuncZeroDelivery/LLGo 66530 ns/op +2253 ns/op / +3.5% (worse)
macOS CreateStop/Go 136.200 ns/op +1.5 ns/op / +1.1% (worse)
macOS CreateStop/LLGo 440.900 ns/op +15.6 ns/op / +3.7% (worse)
macOS RearmStopped/Go 50.990 ns/op -6.37 ns/op / -11.1% (better)
macOS RearmStopped/LLGo 328.800 ns/op -14.9 ns/op / -4.3% (better)
macOS ResetActive/Go 40.980 ns/op -2.89 ns/op / -6.6% (better)
macOS ResetActive/LLGo 155.900 ns/op -10.1 ns/op / -6.1% (better)
macOS ResetHeap1024/Go 37.600 ns/op -6.05 ns/op / -13.9% (better)
macOS ResetHeap1024/LLGo 76.650 ns/op -9.41 ns/op / -10.9% (better)
Windows MinGW AfterFuncZeroDelivery/Go 497.100 ns/op +13.4 ns/op / +2.8% (worse)
Windows MinGW AfterFuncZeroDelivery/LLGo 144296 ns/op +2289 ns/op / +1.6% (worse)
Windows MinGW CreateStop/Go 121.700 ns/op +4.7 ns/op / +4.0% (worse)
Windows MinGW CreateStop/LLGo 484.600 ns/op +28.3 ns/op / +6.2% (worse)
Windows MinGW RearmStopped/Go 31.600 ns/op +0.07 ns/op / +0.2% (worse)
Windows MinGW RearmStopped/LLGo 292.500 ns/op +12.1 ns/op / +4.3% (worse)
Windows MinGW ResetActive/Go 19.070 ns/op -0.07 ns/op / -0.4% (better)
Windows MinGW ResetActive/LLGo 161.300 ns/op +1.1 ns/op / +0.7% (worse)
Windows MinGW ResetHeap1024/Go 19.160 ns/op -0.2 ns/op / -1.0% (better)
Windows MinGW ResetHeap1024/LLGo 141.900 ns/op +3.9 ns/op / +2.8% (worse)
Windows MinGW 386 AfterFuncZeroDelivery/Go 958.200 ns/op -4.6 ns/op / -0.5% (better)
Windows MinGW 386 AfterFuncZeroDelivery/LLGo 201753 ns/op -4245 ns/op / -2.1% (better)
Windows MinGW 386 CreateStop/Go 188.700 ns/op -8.9 ns/op / -4.5% (better)
Windows MinGW 386 CreateStop/LLGo 610.200 ns/op +96.6 ns/op / +18.8% (worse)
Windows MinGW 386 RearmStopped/Go 63.350 ns/op -0.43 ns/op / -0.7% (better)
Windows MinGW 386 RearmStopped/LLGo 370.300 ns/op 0 ns/op / +0.0%
Windows MinGW 386 ResetActive/Go 39.070 ns/op -0.05 ns/op / -0.1% (better)
Windows MinGW 386 ResetActive/LLGo 386.600 ns/op -619.4 ns/op / -61.6% (better)
Windows MinGW 386 ResetHeap1024/Go 39.490 ns/op +0.05 ns/op / +0.1% (worse)
Windows MinGW 386 ResetHeap1024/LLGo 187.500 ns/op -0.3 ns/op / -0.2% (better)
Windows MinGW ARM64 AfterFuncZeroDelivery/Go 669.100 ns/op -2.4 ns/op / -0.4% (better)
Windows MinGW ARM64 AfterFuncZeroDelivery/LLGo 148436 ns/op +1656 ns/op / +1.1% (worse)
Windows MinGW ARM64 CreateStop/Go 200.100 ns/op +5.6 ns/op / +2.9% (worse)
Windows MinGW ARM64 CreateStop/LLGo 362.700 ns/op -13.2 ns/op / -3.5% (better)
Windows MinGW ARM64 RearmStopped/Go 70.600 ns/op +0.16 ns/op / +0.2% (worse)
Windows MinGW ARM64 RearmStopped/LLGo 248.900 ns/op +0.3 ns/op / +0.1% (worse)
Windows MinGW ARM64 ResetActive/Go 31.090 ns/op +0.03 ns/op / +0.1% (worse)
Windows MinGW ARM64 ResetActive/LLGo 118.900 ns/op -9 ns/op / -7.0% (better)
Windows MinGW ARM64 ResetHeap1024/Go 31.220 ns/op +0.25 ns/op / +0.8% (worse)
Windows MinGW ARM64 ResetHeap1024/LLGo 125.100 ns/op +0.1 ns/op / +0.1% (worse)
Windows MSVC AfterFuncZeroDelivery/Go 557.500 ns/op -9.1 ns/op / -1.6% (better)
Windows MSVC AfterFuncZeroDelivery/LLGo 175564 ns/op -16018 ns/op / -8.4% (better)
Windows MSVC CreateStop/Go 114.500 ns/op -1 ns/op / -0.9% (better)
Windows MSVC CreateStop/LLGo 433.300 ns/op +2.4 ns/op / +0.6% (worse)
Windows MSVC RearmStopped/Go 31.510 ns/op +0.15 ns/op / +0.5% (worse)
Windows MSVC RearmStopped/LLGo 345 ns/op +72.2 ns/op / +26.5% (worse)
Windows MSVC ResetActive/Go 20.130 ns/op +0.01 ns/op / +0.0497% (worse)
Windows MSVC ResetActive/LLGo 155.300 ns/op +1.4 ns/op / +0.9% (worse)
Windows MSVC ResetHeap1024/Go 20.440 ns/op -0.01 ns/op / -0.0489% (better)
Windows MSVC ResetHeap1024/LLGo 124.700 ns/op +2 ns/op / +1.6% (worse)
Windows MSVC 386 AfterFuncZeroDelivery/Go 762.500 ns/op -17.5 ns/op / -2.2% (better)
Windows MSVC 386 AfterFuncZeroDelivery/LLGo 134914 ns/op +1407 ns/op / +1.1% (worse)
Windows MSVC 386 CreateStop/Go 170.600 ns/op +3.1 ns/op / +1.9% (worse)
Windows MSVC 386 CreateStop/LLGo 406.200 ns/op +35 ns/op / +9.4% (worse)
Windows MSVC 386 RearmStopped/Go 56.510 ns/op -0.26 ns/op / -0.5% (better)
Windows MSVC 386 RearmStopped/LLGo 265.400 ns/op +5.7 ns/op / +2.2% (worse)
Windows MSVC 386 ResetActive/Go 32.600 ns/op -0.04 ns/op / -0.1% (better)
Windows MSVC 386 ResetActive/LLGo 855.400 ns/op +43 ns/op / +5.3% (worse)
Windows MSVC 386 ResetHeap1024/Go 32.800 ns/op +0.01 ns/op / +0.0305% (worse)
Windows MSVC 386 ResetHeap1024/LLGo 142.300 ns/op -0.1 ns/op / -0.1% (better)
Windows MSVC ARM64 AfterFuncZeroDelivery/Go 668.100 ns/op -0.6 ns/op / -0.1% (better)
Windows MSVC ARM64 AfterFuncZeroDelivery/LLGo 168354 ns/op -197 ns/op / -0.1% (better)
Windows MSVC ARM64 CreateStop/Go 212.100 ns/op +10.8 ns/op / +5.4% (worse)
Windows MSVC ARM64 CreateStop/LLGo 390.100 ns/op +4.7 ns/op / +1.2% (worse)
Windows MSVC ARM64 RearmStopped/Go 70.590 ns/op +0.01 ns/op / +0.01417% (worse)
Windows MSVC ARM64 RearmStopped/LLGo 271.500 ns/op +0.9 ns/op / +0.3% (worse)
Windows MSVC ARM64 ResetActive/Go 31.050 ns/op +0.09 ns/op / +0.3% (worse)
Windows MSVC ARM64 ResetActive/LLGo 133.100 ns/op -3.2 ns/op / -2.3% (better)
Windows MSVC ARM64 ResetHeap1024/Go 31.170 ns/op +0.12 ns/op / +0.4% (worse)
Windows MSVC ARM64 ResetHeap1024/LLGo 136.800 ns/op -0.3 ns/op / -0.2% (better)

Compared with ec9c2488bc37 measured in the same runner job.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant