Skip to content

runtime: fix native and Wasm compatibility regressions - #2751

Merged
cpunion merged 19 commits into
xgo-dev:mainfrom
cpunion:codex/compat-expanded-platform-fixes-20261007
Oct 8, 2026
Merged

cpunion merged 19 commits into
xgo-dev:mainfrom
cpunion:codex/compat-expanded-platform-fixes-20261007

Conversation

@cpunion

@cpunion cpunion commented Oct 7, 2026 •

Copy link
Copy Markdown
Collaborator

Fix incorrect results, crashes, and excessive allocation exposed by native and Wasm compatibility tests.

  • Correct ARM64 C closure environments and GNU fault unwinding.
  • Fix caller filenames, function boundaries, panic frames, and deferred caller lines.
  • Install WASI runtime-local state at every Go entry, including goroutines in programs without standard-library imports. Reserve initial heap space after static data instead of fixing total memory at 64 MiB.
  • Preserve process exit across nested JS callbacks and resumed Asyncify fibers, restoring the caller’s stack bounds before propagating the exit. Supply GoJS browser filesystem fallbacks so importing os does not fail during syscall initialization.
  • Remove unused allocations and integer map-key heap temporaries; improve threaded allocator waits and metadata scans. Keep suspendable finalizers instrumented.

Add executable regressions to existing native/Wasm jobs, including callback exit statuses, worker panic/recover with and without GC, caller lines, and allocation counts. Fix WASI GOROOT filesystem preopens and dependency mirror setup; retain complete stress tests with measured, bounded resource budgets.

Validation: compiler tests pass; the original JS timer/callback case and WASI concurrent panic case each passed 50 runs; worker startup/recovery passed 10 runs in each GC mode; 54 JS host-boundary cases passed across GoJS, Emscripten32, and Emscripten64; optimized Node/Chrome acceptance passed, including browser filesystem initialization. Full external qualification runs in llgo-compat/ci #37.

@fennoai fennoai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed the full diff with four review passes (quality, performance, security, documentation). This is a tightly scoped, well-executed PR; all three fixes are minimal and correct, each backed by a targeted regression test.

Verified correct:

  • AArch64 ELF bridge (libffi.c): the naked trampoline ldr x16,[x18]; ldr x15,[x18,#8]; br x16 correctly moves libffi's X18-delivered static chain into LLVM's X15 nest register before a tail branch. The stack-allocated llgo_ffi_call_context is live for the synchronous ffi_call_go call; X16 (IP0) is corruptible scratch and X15 is the intended nest reg, so argument registers, LR, and SP pass through untouched. The X18-source assumption holds on the guarded ELF targets (not Apple/Android/Windows), matching libffi's own Go-closure ABI. The #error guard still enforces exactly one final-hop path.
  • Caller cache invalidation (caller.go): clearing frame.captured before the field overwrites is correct, and the added || frame.File != file is exactly what fixes the line-directive-with-same-line case. The compare is short-circuited behind frame.Line == line and runs at most once per call, so the hot-path cost is a bounded, allocation-free constant factor — an acceptable trade for the correctness fix.
  • WASI PATH injection (wasm_profile_test.go): args[1:] correctly drops only the leading "run", the []string{"run", "--env=PATH="+path} literal has len==cap so the subsequent append reallocates (no aliasing), and the unit test's expected args confirm the ordering. Values go through exec.Command with no shell — no host injection.

Only minor comment/wording nits below; nothing blocking.

Comment thread runtime/internal/clite/ffi/_wrap/libffi.c Outdated
Comment thread runtime/internal/runtime/caller.go Outdated
@codecov

codecov Bot commented Oct 7, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@github-actions

github-actions Bot commented Oct 7, 2026 •

Copy link
Copy Markdown

LLGo baseline benchmarks

f071b9a2fab2 | workflow run | long-term charts

Program measurements

Platform Workload File size vs base Text size vs base Build vs base Run vs base
Linux cprintf 8568 B 0 B / +0.0% 387 B 0 B / +0.0% 1.039 s +4.967 ms / +0.5% (worse) 1.289 ms +52.28 us / +4.2% (worse)
Linux cprintf-lto 8408 B 0 B / +0.0% 368 B 0 B / +0.0% 1.058 s -33.56 ms / -3.1% (better) 1.280 ms +51.74 us / +4.2% (worse)
Linux fmtprintf 4887240 B +952 B / +0.01948% (worse) 504614 B -147 B / -0.02912% (better) 7.844 s -176.7 ms / -2.2% (better) 3.498 ms +356.2 us / +11.3% (worse)
Linux fmtprintf-lto 3607696 B -8872 B / -0.2% (better) 442924 B -151 B / -0.03408% (better) 17.595 s -804.1 ms / -4.4% (better) 4.547 ms +1.558 ms / +52.1% (worse)
Linux println 681656 B +6616 B / +1.0% (worse) 16855 B 0 B / +0.0% 1.031 s -300.7 ms / -22.6% (better) 1.568 ms -102.3 us / -6.1% (better)
Linux println-lto 190160 B +120 B / +0.1% (worse) 14273 B 0 B / +0.0% 1.380 s -173.5 ms / -11.2% (better) 1.694 ms -241.6 us / -12.5% (better)
macOS cprintf 50736 B 0 B / +0.0% 4409 B 0 B / +0.0% 1.587 s +262.9 ms / +19.9% (worse) 6.386 ms +2.07 ms / +47.9% (worse)
macOS cprintf-lto 50496 B 0 B / +0.0% 161 B 0 B / +0.0% 1.293 s +68.01 ms / +5.6% (worse) 3.227 ms -2.746 ms / -46.0% (better)
macOS fmtprintf 1774368 B 0 B / +0.0% 880617 B +88 B / +0.009994% (worse) 7.931 s +936.9 ms / +13.4% (worse) 11.358 ms -211.5 us / -1.8% (better)
macOS fmtprintf-lto 1361232 B 0 B / +0.0% 764945 B -12 B / -0.001569% (better) 14.096 s -1.373 s / -8.9% (better) 7.074 ms +424.3 us / +6.4% (worse)
macOS println 99344 B 0 B / +0.0% 24216 B 0 B / +0.0% 1.107 s -526.1 ms / -32.2% (better) 3.916 ms -2.516 ms / -39.1% (better)
macOS println-lto 83664 B 0 B / +0.0% 21457 B 0 B / +0.0% 1.304 s -565.4 ms / -30.2% (better) 4.505 ms -1.777 ms / -28.3% (better)
Windows MinGW cprintf 655872 B +4608 B / +0.7% (worse) 4550 B 0 B / +0.0% 1.793 s -115.6 ms / -6.1% (better) 4.356 ms -133.7 us / -3.0% (better)
Windows MinGW cprintf-lto 44032 B 0 B / +0.0% 4502 B 0 B / +0.0% 1.793 s -24.73 ms / -1.4% (better) 4.067 ms -278.5 us / -6.4% (better)
Windows MinGW fmtprintf 5444608 B +4096 B / +0.1% (worse) 604550 B +80 B / +0.01323% (worse) 7.678 s +104.2 ms / +1.4% (worse) 9.048 ms -891.9 us / -9.0% (better)
Windows MinGW fmtprintf-lto 4145664 B +1536 B / +0.03706% (worse) 551670 B +80 B / +0.0145% (worse) 15.583 s +77.59 ms / +0.5% (worse) 9.749 ms +13.2 us / +0.1% (worse)
Windows MinGW println 710656 B +4608 B / +0.7% (worse) 25190 B 0 B / +0.0% 1.766 s -26.58 ms / -1.5% (better) 8.273 ms +330.1 us / +4.2% (worse)
Windows MinGW println-lto 208896 B 0 B / +0.0% 22070 B 0 B / +0.0% 2.059 s -37.24 ms / -1.8% (better) 7.476 ms -478.4 us / -6.0% (better)
Windows MinGW 386 cprintf 605184 B +3584 B / +0.6% (worse) 5352 B 0 B / +0.0% 1.840 s +55.56 ms / +3.1% (worse) 6.418 ms +896.4 us / +16.2% (worse)
Windows MinGW 386 cprintf-lto 103424 B 0 B / +0.0% 5112 B 0 B / +0.0% 1.932 s +115.2 ms / +6.3% (worse) 5.823 ms +85.1 us / +1.5% (worse)
Windows MinGW 386 fmtprintf 4749824 B +4096 B / +0.1% (worse) 472696 B +224 B / +0.04741% (worse) 8.363 s +735.8 ms / +9.6% (worse) 12.410 ms +2.002 ms / +19.2% (worse)
Windows MinGW 386 fmtprintf-lto 4149248 B +512 B / +0.01234% (worse) 451064 B -68 B / -0.01507% (better) 16.177 s +814.9 ms / +5.3% (worse) 12.577 ms -651.8 us / -4.9% (better)
Windows MinGW 386 println 657408 B +3072 B / +0.5% (worse) 21500 B 0 B / +0.0% 1.930 s +114.9 ms / +6.3% (worse) 11.397 ms +902.8 us / +8.6% (worse)
Windows MinGW 386 println-lto 259072 B +512 B / +0.2% (worse) 19328 B 0 B / +0.0% 2.222 s +31.51 ms / +1.4% (worse) 10.302 ms -232.2 us / -2.2% (better)
Windows MinGW ARM64 cprintf 666112 B +4096 B / +0.6% (worse) 4436 B 0 B / +0.0% 2.008 s -11.46 ms / -0.6% (better) 7.054 ms -325.8 us / -4.4% (better)
Windows MinGW ARM64 cprintf-lto 43520 B 0 B / +0.0% 4388 B 0 B / +0.0% 1.981 s -70.39 ms / -3.4% (better) 6.987 ms -90.6 us / -1.3% (better)
Windows MinGW ARM64 fmtprintf 5353984 B +5632 B / +0.1% (worse) 512720 B +56 B / +0.01092% (worse) 7.134 s -52.62 ms / -0.7% (better) 13.457 ms +195.2 us / +1.5% (worse)
Windows MinGW ARM64 fmtprintf-lto 4312576 B +1024 B / +0.02375% (worse) 478928 B +76 B / +0.01587% (worse) 13.802 s -199.1 ms / -1.4% (better) 13.243 ms +265.3 us / +2.0% (worse)
Windows MinGW ARM64 println 717824 B +4096 B / +0.6% (worse) 23912 B 0 B / +0.0% 1.973 s -68.94 ms / -3.4% (better) 11.775 ms -785.6 us / -6.3% (better)
Windows MinGW ARM64 println-lto 215552 B 0 B / +0.0% 21272 B 0 B / +0.0% 2.250 s -41.58 ms / -1.8% (better) 11.339 ms -868.2 us / -7.1% (better)
Windows MSVC cprintf 900096 B +4608 B / +0.5% (worse) 66019 B 0 B / +0.0% 1.921 s +102.4 ms / +5.6% (worse) 3.827 ms -388 us / -9.2% (better)
Windows MSVC cprintf-lto 291328 B 0 B / +0.0% 65971 B 0 B / +0.0% 1.661 s +36.12 ms / +2.2% (worse) 3.380 ms -81.2 us / -2.3% (better)
Windows MSVC fmtprintf 5747712 B +5632 B / +0.1% (worse) 700310 B +96 B / +0.01371% (worse) 7.268 s +243.6 ms / +3.5% (worse) 9.006 ms -731.1 us / -7.5% (better)
Windows MSVC fmtprintf-lto 4470784 B +1536 B / +0.03437% (worse) 651398 B +48 B / +0.007369% (worse) 14.019 s +538.7 ms / +4.0% (worse) 9.777 ms +739.1 us / +8.2% (worse)
Windows MSVC println 1023488 B +4608 B / +0.5% (worse) 121062 B 0 B / +0.0% 1.695 s +71.01 ms / +4.4% (worse) 8.060 ms -200.1 us / -2.4% (better)
Windows MSVC println-lto 529920 B 0 B / +0.0% 118598 B 0 B / +0.0% 1.844 s -9.939 ms / -0.5% (better) 7.192 ms -252.1 us / -3.4% (better)
Windows MSVC 386 cprintf 516608 B +2560 B / +0.5% (worse) 3947 B 0 B / +0.0% 1.565 s -12.79 ms / -0.8% (better) 5.736 ms -79.8 us / -1.4% (better)
Windows MSVC 386 cprintf-lto 44032 B 0 B / +0.0% 3901 B 0 B / +0.0% 1.558 s +6.966 ms / +0.4% (worse) 7.884 ms +2.122 ms / +36.8% (worse)
Windows MSVC 386 fmtprintf 4483072 B +3072 B / +0.1% (worse) 456108 B +208 B / +0.04562% (worse) 7.194 s +136.2 ms / +1.9% (worse) 11.718 ms +755.6 us / +6.9% (worse)
Windows MSVC 386 fmtprintf-lto 3900928 B +1024 B / +0.02626% (worse) 426635 B -32 B / -0.0075% (better) 13.912 s +218.6 ms / +1.6% (worse) 12.625 ms +1.004 ms / +8.6% (worse)
Windows MSVC 386 println 571392 B +3072 B / +0.5% (worse) 20356 B 0 B / +0.0% 1.541 s +1.146 ms / +0.1% (worse) 8.997 ms -127.8 us / -1.4% (better)
Windows MSVC 386 println-lto 199168 B 0 B / +0.0% 18517 B 0 B / +0.0% 1.813 s +12.07 ms / +0.7% (worse) 11.498 ms +1.545 ms / +15.5% (worse)
Windows MSVC ARM64 cprintf 666112 B +3584 B / +0.5% (worse) 4240 B 0 B / +0.0% 1.625 s -4.053 ms / -0.2% (better) 6.403 ms -233.1 us / -3.5% (better)
Windows MSVC ARM64 cprintf-lto 48640 B 0 B / +0.0% 4140 B 0 B / +0.0% 1.619 s +71.96 ms / +4.7% (worse) 6.648 ms +112.5 us / +1.7% (worse)
Windows MSVC ARM64 fmtprintf 5349888 B +4096 B / +0.1% (worse) 512676 B +64 B / +0.01249% (worse) 6.679 s +167.9 ms / +2.6% (worse) 18.344 ms +4.992 ms / +37.4% (worse)
Windows MSVC ARM64 fmtprintf-lto 4319744 B +512 B / +0.01185% (worse) 479600 B +64 B / +0.01335% (worse) 12.849 s +217.6 ms / +1.7% (worse) 13.566 ms -244.6 us / -1.8% (better)
Windows MSVC ARM64 println 719872 B +3584 B / +0.5% (worse) 23956 B 0 B / +0.0% 1.570 s +38.72 ms / +2.5% (worse) 11.798 ms +643.6 us / +5.8% (worse)
Windows MSVC ARM64 println-lto 220672 B 0 B / +0.0% 21444 B 0 B / +0.0% 1.846 s +87.92 ms / +5.0% (worse) 12.448 ms +903.9 us / +7.8% (worse)
Core language and compiler benchmarks
Platform Benchmark ns/op vs base
Linux BenchmarkLookupPCRandom 14.590 ns/op +0.15 ns/op / +1.0% (worse)
Linux BenchmarkMergeCompilerFlags 191.100 ns/op -14.6 ns/op / -7.1% (better)
Linux BenchmarkMergeLinkerFlags 126.700 ns/op -17.2 ns/op / -12.0% (better)
Linux BenchmarkChannelBuffered 55.320 ns/op +0.4 ns/op / +0.7% (worse)
Linux BenchmarkChannelHandoff 13292 ns/op -836 ns/op / -5.9% (better)
Linux BenchmarkDefer 53.640 ns/op +0.43 ns/op / +0.8% (worse)
Linux BenchmarkDirectCall 1.517 ns/op -0.001 ns/op / -0.1% (better)
Linux BenchmarkGlobalRead 1.553 ns/op +0.003 ns/op / +0.2% (worse)
Linux BenchmarkGlobalWrite 7.779 ns/op +0.014 ns/op / +0.2% (worse)
Linux BenchmarkGoroutine 24300 ns/op -122 ns/op / -0.5% (better)
Linux BenchmarkInterfaceCall 6.601 ns/op +0.385 ns/op / +6.2% (worse)
Linux BenchmarkRuntimeGetG 2.529 ns/op -0.409 ns/op / -13.9% (better)
macOS BenchmarkLookupPCRandom 13.030 ns/op -5.21 ns/op / -28.6% (better)
macOS BenchmarkMergeCompilerFlags 124.300 ns/op -42.6 ns/op / -25.5% (better)
macOS BenchmarkMergeLinkerFlags 67.620 ns/op -37.88 ns/op / -35.9% (better)
macOS BenchmarkChannelBuffered 31.300 ns/op +5.95 ns/op / +23.5% (worse)
macOS BenchmarkChannelHandoff 11312 ns/op +3100 ns/op / +37.7% (worse)
macOS BenchmarkDefer 43.290 ns/op -12.13 ns/op / -21.9% (better)
macOS BenchmarkDirectCall 1.221 ns/op +0.005 ns/op / +0.4% (worse)
macOS BenchmarkGlobalRead 1.259 ns/op -0.017 ns/op / -1.3% (better)
macOS BenchmarkGlobalWrite 1.452 ns/op +0.03 ns/op / +2.1% (worse)
macOS BenchmarkGoroutine 68483 ns/op -10125 ns/op / -12.9% (better)
macOS BenchmarkInterfaceCall 4.209 ns/op -1.39 ns/op / -24.8% (better)
macOS BenchmarkRuntimeGetG 2.425 ns/op -0.095 ns/op / -3.8% (better)
Windows MinGW BenchmarkLookupPCRandom 11.800 ns/op -0.08 ns/op / -0.7% (better)
Windows MinGW BenchmarkMergeCompilerFlags 588.300 ns/op -10.4 ns/op / -1.7% (better)
Windows MinGW BenchmarkMergeLinkerFlags 512.400 ns/op -3.6 ns/op / -0.7% (better)
Windows MinGW BenchmarkChannelBuffered 45.090 ns/op -0.1 ns/op / -0.2% (better)
Windows MinGW BenchmarkChannelHandoff 1550 ns/op +458 ns/op / +41.9% (worse)
Windows MinGW BenchmarkDefer 70.970 ns/op -0.01 ns/op / -0.01409% (better)
Windows MinGW BenchmarkDirectCall 1.181 ns/op +0.038 ns/op / +3.3% (worse)
Windows MinGW BenchmarkGlobalRead 1.053 ns/op +0.023 ns/op / +2.2% (worse)
Windows MinGW BenchmarkGlobalWrite 8.710 ns/op -0.03 ns/op / -0.3% (better)
Windows MinGW BenchmarkGoroutine 83696 ns/op -2926 ns/op / -3.4% (better)
Windows MinGW BenchmarkInterfaceCall 5.889 ns/op +0.306 ns/op / +5.5% (worse)
Windows MinGW BenchmarkRuntimeGetG 2.155 ns/op +0.141 ns/op / +7.0% (worse)
Windows MinGW 386 BenchmarkLookupPCRandom 26.760 ns/op -0.03 ns/op / -0.1% (better)
Windows MinGW 386 BenchmarkMergeCompilerFlags 764.400 ns/op -64.6 ns/op / -7.8% (better)
Windows MinGW 386 BenchmarkMergeLinkerFlags 712.800 ns/op -29.3 ns/op / -3.9% (better)
Windows MinGW 386 BenchmarkChannelBuffered 39.280 ns/op -0.62 ns/op / -1.6% (better)
Windows MinGW 386 BenchmarkChannelHandoff 828.300 ns/op -89.6 ns/op / -9.8% (better)
Windows MinGW 386 BenchmarkDefer 45.990 ns/op -0.46 ns/op / -1.0% (better)
Windows MinGW 386 BenchmarkDirectCall 1.550 ns/op -0.004 ns/op / -0.3% (better)
Windows MinGW 386 BenchmarkGlobalRead 1.861 ns/op -0.002 ns/op / -0.1% (better)
Windows MinGW 386 BenchmarkGlobalWrite 7.776 ns/op -0.014 ns/op / -0.2% (better)
Windows MinGW 386 BenchmarkGoroutine 106454 ns/op -4837 ns/op / -4.3% (better)
Windows MinGW 386 BenchmarkInterfaceCall 8.186 ns/op +0.039 ns/op / +0.5% (worse)
Windows MinGW 386 BenchmarkRuntimeGetG 2.489 ns/op +0.311 ns/op / +14.3% (worse)
Windows MinGW ARM64 BenchmarkLookupPCRandom 12.090 ns/op +0.01 ns/op / +0.1% (worse)
Windows MinGW ARM64 BenchmarkMergeCompilerFlags 569.100 ns/op +3.7 ns/op / +0.7% (worse)
Windows MinGW ARM64 BenchmarkMergeLinkerFlags 525.100 ns/op -8.9 ns/op / -1.7% (better)
Windows MinGW ARM64 BenchmarkChannelBuffered 37.650 ns/op -1.13 ns/op / -2.9% (better)
Windows MinGW ARM64 BenchmarkChannelHandoff 2308 ns/op -163 ns/op / -6.6% (better)
Windows MinGW ARM64 BenchmarkDefer 55.870 ns/op -0.02 ns/op / -0.03578% (better)
Windows MinGW ARM64 BenchmarkDirectCall 0.664 ns/op +0.0005 ns/op / +0.1% (worse)
Windows MinGW ARM64 BenchmarkGlobalRead 0.590 ns/op +0.0002 ns/op / +0.03392% (worse)
Windows MinGW ARM64 BenchmarkGlobalWrite 0.737 ns/op +0.0005 ns/op / +0.1% (worse)
Windows MinGW ARM64 BenchmarkGoroutine 63997 ns/op -1582 ns/op / -2.4% (better)
Windows MinGW ARM64 BenchmarkInterfaceCall 4.149 ns/op +0.001 ns/op / +0.02411% (worse)
Windows MinGW ARM64 BenchmarkRuntimeGetG 1.800 ns/op +0.029 ns/op / +1.6% (worse)
Windows MSVC BenchmarkLookupPCRandom 13.160 ns/op +0.15 ns/op / +1.2% (worse)
Windows MSVC BenchmarkMergeCompilerFlags 603.800 ns/op -17.6 ns/op / -2.8% (better)
Windows MSVC BenchmarkMergeLinkerFlags 532.800 ns/op -11.9 ns/op / -2.2% (better)
Windows MSVC BenchmarkChannelBuffered 28.380 ns/op -0.54 ns/op / -1.9% (better)
Windows MSVC BenchmarkChannelHandoff 1077 ns/op +22 ns/op / +2.1% (worse)
Windows MSVC BenchmarkDefer 54.360 ns/op -2.61 ns/op / -4.6% (better)
Windows MSVC BenchmarkDirectCall 1.546 ns/op -0.002 ns/op / -0.1% (better)
Windows MSVC BenchmarkGlobalRead 1.549 ns/op -0.001 ns/op / -0.1% (better)
Windows MSVC BenchmarkGlobalWrite 2.472 ns/op -0.005 ns/op / -0.2% (better)
Windows MSVC BenchmarkGoroutine 89170 ns/op -1178 ns/op / -1.3% (better)
Windows MSVC BenchmarkInterfaceCall 8.376 ns/op -0.01 ns/op / -0.1% (better)
Windows MSVC BenchmarkRuntimeGetG 2.169 ns/op 0 ns/op / +0.0%
Windows MSVC 386 BenchmarkLookupPCRandom 26.650 ns/op +0.08 ns/op / +0.3% (worse)
Windows MSVC 386 BenchmarkMergeCompilerFlags 794.600 ns/op +19.4 ns/op / +2.5% (worse)
Windows MSVC 386 BenchmarkMergeLinkerFlags 716.700 ns/op +5.7 ns/op / +0.8% (worse)
Windows MSVC 386 BenchmarkChannelBuffered 39.190 ns/op -0.01 ns/op / -0.02551% (better)
Windows MSVC 386 BenchmarkChannelHandoff 820.300 ns/op -44.9 ns/op / -5.2% (better)
Windows MSVC 386 BenchmarkDefer 47.590 ns/op -0.73 ns/op / -1.5% (better)
Windows MSVC 386 BenchmarkDirectCall 1.556 ns/op +0.005 ns/op / +0.3% (worse)
Windows MSVC 386 BenchmarkGlobalRead 1.550 ns/op 0 ns/op / +0.0%
Windows MSVC 386 BenchmarkGlobalWrite 7.800 ns/op +0.008 ns/op / +0.1% (worse)
Windows MSVC 386 BenchmarkGoroutine 115239 ns/op +3376 ns/op / +3.0% (worse)
Windows MSVC 386 BenchmarkInterfaceCall 8.411 ns/op +0.294 ns/op / +3.6% (worse)
Windows MSVC 386 BenchmarkRuntimeGetG 2.169 ns/op +0.236 ns/op / +12.2% (worse)
Windows MSVC ARM64 BenchmarkLookupPCRandom 12.040 ns/op +0.07 ns/op / +0.6% (worse)
Windows MSVC ARM64 BenchmarkMergeCompilerFlags 557.500 ns/op -7.5 ns/op / -1.3% (better)
Windows MSVC ARM64 BenchmarkMergeLinkerFlags 533.800 ns/op +7.6 ns/op / +1.4% (worse)
Windows MSVC ARM64 BenchmarkChannelBuffered 38.860 ns/op +1.25 ns/op / +3.3% (worse)
Windows MSVC ARM64 BenchmarkChannelHandoff 2549 ns/op +246 ns/op / +10.7% (worse)
Windows MSVC ARM64 BenchmarkDefer 59.270 ns/op -2.27 ns/op / -3.7% (better)
Windows MSVC ARM64 BenchmarkDirectCall 0.590 ns/op +0.0001 ns/op / +0.01697% (worse)
Windows MSVC ARM64 BenchmarkGlobalRead 0.590 ns/op -0.0001 ns/op / -0.01696% (better)
Windows MSVC ARM64 BenchmarkGlobalWrite 3.832 ns/op +0.002 ns/op / +0.1% (worse)
Windows MSVC ARM64 BenchmarkGoroutine 58586 ns/op +3043 ns/op / +5.5% (worse)
Windows MSVC ARM64 BenchmarkInterfaceCall 4.151 ns/op +0.012 ns/op / +0.3% (worse)
Windows MSVC ARM64 BenchmarkRuntimeGetG 1.769 ns/op -0.035 ns/op / -1.9% (better)
Timer runtime benchmarks
Platform Operation and runtime ns/op vs base
Linux AfterFuncZeroDelivery/Go 903.900 ns/op -18.1 ns/op / -2.0% (better)
Linux AfterFuncZeroDelivery/LLGo 38852 ns/op +674 ns/op / +1.8% (worse)
Linux CreateStop/Go 319.800 ns/op +27.8 ns/op / +9.5% (worse)
Linux CreateStop/LLGo 1583 ns/op -330 ns/op / -17.3% (better)
Linux RearmStopped/Go 128.100 ns/op +11.9 ns/op / +10.2% (worse)
Linux RearmStopped/LLGo 1124 ns/op -107 ns/op / -8.7% (better)
Linux ResetActive/Go 69.680 ns/op +0.97 ns/op / +1.4% (worse)
Linux ResetActive/LLGo 672.700 ns/op -62.3 ns/op / -8.5% (better)
Linux ResetHeap1024/Go 67.160 ns/op -0.01 ns/op / -0.01489% (better)
Linux ResetHeap1024/LLGo 163.100 ns/op -17 ns/op / -9.4% (better)
macOS AfterFuncZeroDelivery/Go 602.100 ns/op -6.6 ns/op / -1.1% (better)
macOS AfterFuncZeroDelivery/LLGo 121596 ns/op +19883 ns/op / +19.5% (worse)
macOS CreateStop/Go 150.800 ns/op -84.4 ns/op / -35.9% (better)
macOS CreateStop/LLGo 633.700 ns/op -324.5 ns/op / -33.9% (better)
macOS RearmStopped/Go 73.210 ns/op +1.06 ns/op / +1.5% (worse)
macOS RearmStopped/LLGo 513.200 ns/op -200.5 ns/op / -28.1% (better)
macOS ResetActive/Go 53.170 ns/op -2.05 ns/op / -3.7% (better)
macOS ResetActive/LLGo 227.800 ns/op +55.5 ns/op / +32.2% (worse)
macOS ResetHeap1024/Go 55.780 ns/op +2.68 ns/op / +5.0% (worse)
macOS ResetHeap1024/LLGo 153.800 ns/op +33.9 ns/op / +28.3% (worse)
Windows MinGW AfterFuncZeroDelivery/Go 724.900 ns/op +9 ns/op / +1.3% (worse)
Windows MinGW AfterFuncZeroDelivery/LLGo 158426 ns/op +6841 ns/op / +4.5% (worse)
Windows MinGW CreateStop/Go 193.500 ns/op -4.8 ns/op / -2.4% (better)
Windows MinGW CreateStop/LLGo 725.400 ns/op -148.1 ns/op / -17.0% (better)
Windows MinGW RearmStopped/Go 71.760 ns/op -0.55 ns/op / -0.8% (better)
Windows MinGW RearmStopped/LLGo 287.700 ns/op -51.9 ns/op / -15.3% (better)
Windows MinGW ResetActive/Go 31.950 ns/op -0.1 ns/op / -0.3% (better)
Windows MinGW ResetActive/LLGo 177.300 ns/op -36.7 ns/op / -17.1% (better)
Windows MinGW ResetHeap1024/Go 32.180 ns/op -0.26 ns/op / -0.8% (better)
Windows MinGW ResetHeap1024/LLGo 118.300 ns/op -16.1 ns/op / -12.0% (better)
Windows MinGW 386 AfterFuncZeroDelivery/Go 959.700 ns/op -8.3 ns/op / -0.9% (better)
Windows MinGW 386 AfterFuncZeroDelivery/LLGo 200536 ns/op -1871 ns/op / -0.9% (better)
Windows MinGW 386 CreateStop/Go 196.300 ns/op -11.9 ns/op / -5.7% (better)
Windows MinGW 386 CreateStop/LLGo 469.500 ns/op -73.3 ns/op / -13.5% (better)
Windows MinGW 386 RearmStopped/Go 63.540 ns/op -0.44 ns/op / -0.7% (better)
Windows MinGW 386 RearmStopped/LLGo 317.900 ns/op -38.4 ns/op / -10.8% (better)
Windows MinGW 386 ResetActive/Go 39.150 ns/op -0.16 ns/op / -0.4% (better)
Windows MinGW 386 ResetActive/LLGo 931.200 ns/op -34.9 ns/op / -3.6% (better)
Windows MinGW 386 ResetHeap1024/Go 39.550 ns/op -0.12 ns/op / -0.3% (better)
Windows MinGW 386 ResetHeap1024/LLGo 175 ns/op -12.8 ns/op / -6.8% (better)
Windows MinGW ARM64 AfterFuncZeroDelivery/Go 667.700 ns/op +6 ns/op / +0.9% (worse)
Windows MinGW ARM64 AfterFuncZeroDelivery/LLGo 146507 ns/op -34 ns/op / -0.0232% (better)
Windows MinGW ARM64 CreateStop/Go 199.300 ns/op +1.6 ns/op / +0.8% (worse)
Windows MinGW ARM64 CreateStop/LLGo 324.600 ns/op -39.3 ns/op / -10.8% (better)
Windows MinGW ARM64 RearmStopped/Go 70.570 ns/op -0.01 ns/op / -0.01417% (better)
Windows MinGW ARM64 RearmStopped/LLGo 210.500 ns/op -41.7 ns/op / -16.5% (better)
Windows MinGW ARM64 ResetActive/Go 30.930 ns/op -0.19 ns/op / -0.6% (better)
Windows MinGW ARM64 ResetActive/LLGo 107.400 ns/op -18.5 ns/op / -14.7% (better)
Windows MinGW ARM64 ResetHeap1024/Go 31.080 ns/op -0.06 ns/op / -0.2% (better)
Windows MinGW ARM64 ResetHeap1024/LLGo 110.400 ns/op -14.7 ns/op / -11.8% (better)
Windows MSVC AfterFuncZeroDelivery/Go 552.600 ns/op -14.7 ns/op / -2.6% (better)
Windows MSVC AfterFuncZeroDelivery/LLGo 176699 ns/op -1120 ns/op / -0.6% (better)
Windows MSVC CreateStop/Go 115.100 ns/op -0.9 ns/op / -0.8% (better)
Windows MSVC CreateStop/LLGo 391.500 ns/op -24.6 ns/op / -5.9% (better)
Windows MSVC RearmStopped/Go 31.390 ns/op -0.02 ns/op / -0.1% (better)
Windows MSVC RearmStopped/LLGo 224.400 ns/op -25.6 ns/op / -10.2% (better)
Windows MSVC ResetActive/Go 20.040 ns/op +0.01 ns/op / +0.04993% (worse)
Windows MSVC ResetActive/LLGo 151.900 ns/op +3.5 ns/op / +2.4% (worse)
Windows MSVC ResetHeap1024/Go 20.420 ns/op -0.14 ns/op / -0.7% (better)
Windows MSVC ResetHeap1024/LLGo 119 ns/op -5.6 ns/op / -4.5% (better)
Windows MSVC 386 AfterFuncZeroDelivery/Go 975.700 ns/op +12.6 ns/op / +1.3% (worse)
Windows MSVC 386 AfterFuncZeroDelivery/LLGo 206921 ns/op -4741 ns/op / -2.2% (better)
Windows MSVC 386 CreateStop/Go 198.800 ns/op +1.3 ns/op / +0.7% (worse)
Windows MSVC 386 CreateStop/LLGo 443.800 ns/op -45.8 ns/op / -9.4% (better)
Windows MSVC 386 RearmStopped/Go 63.350 ns/op +0.01 ns/op / +0.01579% (worse)
Windows MSVC 386 RearmStopped/LLGo 293.400 ns/op -29.7 ns/op / -9.2% (better)
Windows MSVC 386 ResetActive/Go 38.930 ns/op -0.04 ns/op / -0.1% (better)
Windows MSVC 386 ResetActive/LLGo 898.900 ns/op -14 ns/op / -1.5% (better)
Windows MSVC 386 ResetHeap1024/Go 39.350 ns/op +0.04 ns/op / +0.1% (worse)
Windows MSVC 386 ResetHeap1024/LLGo 160 ns/op -14.2 ns/op / -8.2% (better)
Windows MSVC ARM64 AfterFuncZeroDelivery/Go 660.500 ns/op -3.9 ns/op / -0.6% (better)
Windows MSVC ARM64 AfterFuncZeroDelivery/LLGo 135824 ns/op -4801 ns/op / -3.4% (better)
Windows MSVC ARM64 CreateStop/Go 194.200 ns/op -2.1 ns/op / -1.1% (better)
Windows MSVC ARM64 CreateStop/LLGo 351.100 ns/op -66.1 ns/op / -15.8% (better)
Windows MSVC ARM64 RearmStopped/Go 70.620 ns/op +0.02 ns/op / +0.02833% (worse)
Windows MSVC ARM64 RearmStopped/LLGo 224.200 ns/op -53.8 ns/op / -19.4% (better)
Windows MSVC ARM64 ResetActive/Go 30.930 ns/op -0.17 ns/op / -0.5% (better)
Windows MSVC ARM64 ResetActive/LLGo 106.100 ns/op -26.9 ns/op / -20.2% (better)
Windows MSVC ARM64 ResetHeap1024/Go 31.170 ns/op +0.16 ns/op / +0.5% (worse)
Windows MSVC ARM64 ResetHeap1024/LLGo 117.400 ns/op -19.4 ns/op / -14.2% (better)

Compared with b86d349178d6 measured in the same runner job.

@github-actions

github-actions Bot commented Oct 7, 2026 •

Copy link
Copy Markdown

LLGo WebAssembly build benchmarks

f071b9a2fab2 | workflow run | long-term charts

WebAssembly output sizes
Example, profile and compiler Wasm module vs base Generated JS glue vs base
cprintf/j32-emscripten/LLGo 144132 B -5108 B / -3.4% (better) 88157 B 0 B / +0.0%
cprintf/j32-goos-js/LLGo 143452 B -5059 B / -3.4% (better) 72985 B 0 B / +0.0%
cprintf/j64-emscripten-memory64/LLGo 132190 B -4517 B / -3.3% (better) 91328 B 0 B / +0.0%
cprintf/w32-goos-wasip1/LLGo 155051 B +755 B / +0.5% (worse) 0 B 0 B / 0.0%
cprintf/w32-wasi/LLGo 155729 B +750 B / +0.5% (worse) 0 B 0 B / 0.0%
fmtprintf/j32-emscripten/LLGo 3004231 B -52768 B / -1.7% (better) 130661 B +288 B / +0.2% (worse)
fmtprintf/j32-goos-js/Go 2526852 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/j32-goos-js/LLGo 3099056 B -53923 B / -1.7% (better) 102110 B +1337 B / +1.3% (worse)
fmtprintf/j64-emscripten-memory64/LLGo 2766869 B -52223 B / -1.9% (better) 136141 B +288 B / +0.2% (worse)
fmtprintf/w32-goos-wasip1/Go 2500019 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/w32-goos-wasip1/LLGo 2351753 B -3695 B / -0.2% (better) 0 B 0 B / 0.0%
fmtprintf/w32-wasi/LLGo 2348435 B -3591 B / -0.2% (better) 0 B 0 B / 0.0%
j32-emscripten/LLGo 144835 B -3747 B / -2.5% (better) 88157 B 0 B / +0.0%
j32-goos-js/Go 1895533 B 0 B / +0.0% 0 B 0 B / 0.0%
j32-goos-js/LLGo 144298 B -3704 B / -2.5% (better) 72985 B 0 B / +0.0%
j64-emscripten-memory64/LLGo 132586 B -3461 B / -2.5% (better) 91328 B 0 B / +0.0%
reflectcall/j32-emscripten/LLGo 1448836 B -32287 B / -2.2% (better) 104917 B 0 B / +0.0%
reflectcall/j32-goos-js/Go 2191221 B 0 B / +0.0% 0 B 0 B / 0.0%
reflectcall/j32-goos-js/LLGo 1498569 B -32410 B / -2.1% (better) 89743 B 0 B / +0.0%
reflectcall/j64-emscripten-memory64/LLGo 1341384 B -31864 B / -2.3% (better) 109492 B 0 B / +0.0%
reflectcall/w32-goos-wasip1/Go 2205707 B 0 B / +0.0% 0 B 0 B / 0.0%
reflectcall/w32-goos-wasip1/LLGo 1278952 B -4943 B / -0.4% (better) 0 B 0 B / 0.0%
reflectcall/w32-wasi/LLGo 1276580 B -4840 B / -0.4% (better) 0 B 0 B / 0.0%
w32-goos-wasip1/Go 1909947 B 0 B / +0.0% 0 B 0 B / 0.0%
w32-goos-wasip1/LLGo 154698 B +755 B / +0.5% (worse) 0 B 0 B / 0.0%
w32-wasi/LLGo 155376 B +750 B / +0.5% (worse) 0 B 0 B / 0.0%
LLGo WebAssembly build measurements
Example and profile Build vs base
j32-emscripten 7.261 s +352.8 ms / +5.1% (worse)
j32-goos-js 6.869 s +117.3 ms / +1.7% (worse)
j64-emscripten-memory64 6.175 s +126 ms / +2.1% (worse)
reflectcall/w32-wasi 22.950 s -210.5 ms / -0.9% (better)
w32-goos-wasip1 4.248 s +1.105 ms / +0.02602% (worse)
w32-wasi 4.164 s +172.5 ms / +4.3% (worse)

Compared with b86d349178d6 measured in the same runner job.

@cpunion cpunion changed the title runtime: fix ARM64 closure calls and WASI compatibility runtime: fix native and Wasm compatibility regressions Oct 7, 2026

@visualfc visualfc left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approve. The ARM64 nest bridge, GNU unwind X30 register, Caller/gopanic frames, entry-slack gate, address waits, and occupied-metadata skip all look correct, and the new regressions pin those contracts.

Two P2 follow-ups below. Neither blocks merge.

Comment thread runtime/internal/runtime/map_fast64.go
Comment thread runtime/internal/runtime/_wrap/wasi_gc_world.c
@cpunion
cpunion force-pushed the codex/compat-expanded-platform-fixes-20261007 branch from a9bfaf0 to 53c1207 Compare October 8, 2026 08:50
@cpunion
cpunion merged commit 358e17b into xgo-dev:main Oct 8, 2026
99 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants