Skip to content

fix(wasm): stabilize WAMR thread GC and EH boundaries - #2695

Merged
cpunion merged 7 commits into
xgo-dev:mainfrom
cpunion:codex/wasi-thread-stability
Sep 30, 2026
Merged

cpunion merged 7 commits into
xgo-dev:mainfrom
cpunion:codex/wasi-thread-stability

Conversation

@cpunion

@cpunion cpunion commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator

Caught Wasm exceptions in WAMR's classic interpreter can briefly broadcast a cluster-wide termination signal while unwinding to a Wasm caller. A sibling goroutine can exit during that window, causing deferred Goexit to hang or return without running all expected code.

Apply a WAMR 2.4.5 interpreter patch that unwinds directly to the Wasm caller, plus the POSIX signal-handler backport from wasm-micro-runtime/wasm-micro-runtime#5119. Cache identity includes both patches. Uncaught exceptions escaping the native invocation still terminate execution.

Runtime pthread waits also blocked collection while reacquiring locks held by stopped Go threads. Publish suspended callers' roots around the known mutex/condition waits (including timers), and keep those callers out of Go until GC resumes. Initial Goexit unregisters the thread that will never return to Go. Use the same GC-safe pthread mutex for the allocator instead of a host-polling spin lock; let symbol-table initialization waiters acknowledge GC. Smaller initial arenas and a Release WAMR interpreter reduce collection overhead.

Add repeated GC/nogc regression coverage for cross-function panic/recover, concurrent C setjmp/longjmp, deferred worker Goexit, main/init Goexit, unrecovered panic, and a raw escaping Wasm exception. Add finalizer/reflect GC acceptance and a pthread regression for mutex reacquisition during GC. The reflection type-name lookup avoids temporary string allocation. Record the EH comparison in dev/wasm-eh-comparison.md and failure modes/check commands in dev/wasm-wasi-validation.md.

Rebased onto main at 653957f, including merged #2669, #2539, and #2700. This PR includes the remaining EH validation rather than splitting it into another PR. Fixes #2676.

Validation on macOS arm64:

  • bash dev/build_iwasm.sh builds and installs the patched runner.
  • python3 dev/test_wasm_wasi_threads.py passes, including threaded GC, filesystem, selected standard-library packages, and GOROOT sentinel.
  • The expanded suite also passes finalizer and reflect goroutine-startup tests with concurrent GC, and confirms workers can collect after initial Goexit. The native pthread wait regression and native reflection regression pass.
  • The original deferred-Goexit artifact passes 100/100 repetitions with the fix; stock 2.4.5 passes 40/50 and fails by hanging or exiting early.
  • python3 dev/compare_wasm_eh.py --browser with LLGo Binaryen llgo-v132.3 passes all C++ encoding variants in Node/Chrome, Go panic/recover, and O0/O2 Go/C++ wrappers.
  • The complete WAMR test/go binary passes all 242 top-level tests. Concurrent function-info lookup took 9.47 seconds in that run; the previous allocator exceeded the full suite's 12-minute deadline in that test.
  • Full-audit shard 1/16 passes. Every package other than test/go in shard 9/16 passes, including crypto/elliptic; the subsequent complete test/go run verifies the remaining failure is fixed. This is not a claim that the entire standard library has been re-audited on every host. CI runs Linux acceptance and the cross-platform WAMR builds.

Review follow-up: safepoints use an atomic epoch fast path; known blocked conditions/timers no longer poll every 20 ms. Block-state walks carry their segment, address/block lookup uses binary search, and metadata layout is shared by initial allocation and contiguous growth without clearing live metadata. Regression coverage includes 128 out-of-order arenas and indefinite pthread wait/reacquisition. The GC fixture compiles once; finalizers/callback/symbols, the 20 pointer-argument startup races, and the 20 zero-argument startup races execute in three fresh WAMR invocations, each bounded by 300 seconds. Linux amd64 validation of the unchanged runtime completed those groups in 15.94, 36.81 and 11.51 seconds. Every test and repetition remains enabled. Linux arm64 build and execution of the focused regression both pass locally; CI remains the cross-platform gate.

The HTML host self-build and per-worker callback bridge fixes now come from main via #2700; duplicate contribution commits have been removed. Linux host Go tests and patch coverage passed before this rebase; fresh CI validates the rebased head.

@fennoai fennoai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review: WASI thread stability

Reviewed the full diff (base ef3fce4 → head 0ee628d) across four dimensions (code quality, performance, security, documentation). The threading/GC design is careful and well-documented: a coherent stop-the-world protocol, per-thread root chains, a segmented heap, and the C/Go Context layout coupling guarded by _Static_assert. Build-tag partitioning (default/stub vs. llgo.wasi_threads && llgo.wasm.gc.linear) is consistent with no gaps or overlaps. Test coverage of the new wasm-wasi-threaded-gc fixture is thorough (cross-thread pointer handoff, parked receivers, blocked-in-C collection skipping, arena growth). Documentation changes were verified accurate, and the new PEM fixtures are self-consistent, repository-unique test keys (not reused/compromised production keys).

No blocking correctness defects. The notes below are worth considering; one inline nit and a few design-level observations.

Performance (steady-state cost of the new WASI-threaded-GC build)

  • Per-safepoint mutex on the hottest path. wasiGCSafepoint → llgo_wasi_gc_pending (runtime/internal/runtime/_wrap/wasi_gc_world.c) takes and releases the process-wide world_mutex unconditionally on every function entry / loop back-edge. The pending check only reads world_epoch (a uint32) plus the owner; a lock-free atomic epoch read on the fast path (taking the lock only when a stop appears pending) would remove a large fixed tax and cross-thread cache-line contention on all Go execution in this build.
  • O(segments) linear scans inside O(heap) GC loops. segmentForBlock/segmentForAddress/nextSegmentBlock (tinygogc/segments.go) linear-scan up to 128 segments, and they are called per scanned word in startMark, per block in the Alloc free scan, and per block in sweep. With 32 MiB arena growth this makes mark/sweep/alloc effectively O(heap·segments) for larger heaps. Since segments are contiguous in block space, deriving the segment during the walk (or a sorted array + binary search) would restore near-O(1). Relatedly, gcStateOf resolves the segment twice per call (gcStateByteOf then gcStateFromByte) — passing the resolved segment/state through would halve the scans.

Correctness / behavior

  • Parked initial thread stays registered in the GC world. parkInitialWasiThread (runtime/internal/runtime/goexit_initial_wasi_threads.go) sets the initial G dead and, when not the last goroutine, spins in for { c.Usleep(1000) } without unregistering from the GC world. Parked in uninstrumented C, it never reaches a safepoint, so subsequent llgo_wasi_gc_stop calls hit the 500ms timeout and skip sweeping. The tested wasm-wasi-main-goexit* fixtures all converge to the deadlock exit so this path isn't exercised, but a program with a long-lived allocating worker after main Goexits would see repeated 500ms GC stalls and arena-only growth toward OOM. The surrounding comment documents the intent (keep the environment alive to avoid resuming a dead main), so this is a known tradeoff — worth either unregistering from the GC world before the idle loop or documenting the OOM/stall consequence explicitly.

Minor

  • Inline nit on the redundant boolean clause in segments.go (see inline comment).

Comment thread runtime/internal/runtime/tinygogc/segments.go Outdated
@codecov

codecov Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@cpunion cpunion changed the title fix(wasm): isolate WAMR thread exceptions and validate EH boundaries fix(wasm): stabilize WAMR thread GC and EH boundaries Sep 29, 2026
@cpunion
cpunion force-pushed the codex/wasi-thread-stability branch from 047fcd0 to 9a2cd3c Compare September 29, 2026 04:52
@cpunion

cpunion commented Sep 29, 2026

Copy link
Copy Markdown
Collaborator Author

Addressed the review and rebased onto current main (including merged #2539). The branch remains ready for review.

  • Safepoints first load an atomic epoch; the process-wide mutex is acquired only when a stop may be pending.
  • Segment-aware hot loops and binary-search lookups remove repeated linear scans. Initial setup and contiguous growth share the metadata calculation; growth still preserves existing allocation metadata. The new regression covers all 128 out-of-order arenas.
  • Initial Goexit unregistering was already fixed in the prior update; the acceptance fixture verifies collection advances in the surviving timer worker.
  • Known pthread condition waits now stay blocked until signaled; timer waits use their real deadlines. The 500 ms uncooperative-C deadline has a documented constant.
  • The CI GC probe now separates cold compilation from verbose execution, each bounded by 300 seconds, instead of hiding both inside one captured subprocess. Focused Linux arm64 compilation/execution pass locally; the complete revised macOS WAMR acceptance script passes (GC probe compilation 27.57 s, execution 36.24 s).

Added root-selection coverage for threaded GC versus nogc/single-worker profiles. Fresh CI is queued; no coverage threshold was relaxed.

@cpunion

cpunion commented Sep 29, 2026

Copy link
Copy Markdown
Collaborator Author

Follow-up CI fixes are now pushed in fb2a438 (still ready for review):

  • The GC timeout was in execution, not compilation: compilation took 60 seconds, and the pointer/zero startup races each took roughly 134–160 seconds. Compile once and run the two 20-goroutine startup shapes plus the finalizer/callback/symbol group in separate WAMR invocations, each retaining a 300-second execution deadline. No test or repetition was removed. Linux amd64 validation of the unchanged runtime passed in 15.94 / 36.81 / 11.51 seconds; the complete revised acceptance script also passes on macOS.
  • The merged wasm: Emscripten browser fs host without fs.*Sync overlay #2539 callback-installation guard was shared across workers. Reused the thread-local guard fix already in feat(wasm): integrate browser filesystem with bounded workers #2696. The old code reproduces invoke is not a function; Memory32 and Memory64 two-worker concurrent output/filesystem checks pass with the fix.

The prior head's full Linux Go tests and patch-coverage gate passed. Fresh CI for this script/C-only follow-up is queued. The allocator/yield experiments made while investigating timing did not produce reliable gains and are not included.

@github-actions

github-actions Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

LLGo WebAssembly build benchmarks

f46c1a839521 | workflow run | long-term charts

WebAssembly output sizes
Example, profile and compiler Wasm module vs base Generated JS glue vs base
cprintf/j32-emscripten/LLGo 154193 B +685 B / +0.4% (worse) 74890 B 0 B / +0.0%
cprintf/j32-goos-js/LLGo 152295 B +1299 B / +0.9% (worse) 73165 B 0 B / +0.0%
cprintf/j64-emscripten-memory64/LLGo 141248 B +579 B / +0.4% (worse) 78778 B 0 B / +0.0%
cprintf/w32-goos-wasip1/LLGo 147689 B +993 B / +0.7% (worse) 0 B 0 B / 0.0%
cprintf/w32-wasi/LLGo 147797 B +436 B / +0.3% (worse) 0 B 0 B / 0.0%
fmtprintf/j32-emscripten/LLGo 3213428 B +1851 B / +0.1% (worse) 118590 B 0 B / +0.0%
fmtprintf/j32-goos-js/Go 2526852 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/j32-goos-js/LLGo 3192523 B +2487 B / +0.1% (worse) 101635 B 0 B / +0.0%
fmtprintf/j64-emscripten-memory64/LLGo 2950033 B +1699 B / +0.1% (worse) 125433 B 0 B / +0.0%
fmtprintf/w32-goos-wasip1/Go 2500019 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/w32-goos-wasip1/LLGo 2844343 B +2157 B / +0.1% (worse) 0 B 0 B / 0.0%
fmtprintf/w32-wasi/LLGo 2711131 B +1570 B / +0.1% (worse) 0 B 0 B / 0.0%
j32-emscripten/LLGo 153428 B +730 B / +0.5% (worse) 74890 B 0 B / +0.0%
j32-goos-js/Go 1895533 B 0 B / +0.0% 0 B 0 B / 0.0%
j32-goos-js/LLGo 151768 B +1341 B / +0.9% (worse) 73165 B 0 B / +0.0%
j64-emscripten-memory64/LLGo 140582 B +633 B / +0.5% (worse) 78778 B 0 B / +0.0%
reflectcall/j32-emscripten/LLGo 1543771 B +1697 B / +0.1% (worse) 92056 B 0 B / +0.0%
reflectcall/j32-goos-js/Go 2191221 B 0 B / +0.0% 0 B 0 B / 0.0%
reflectcall/j32-goos-js/LLGo 1546222 B +2328 B / +0.2% (worse) 90331 B 0 B / +0.0%
reflectcall/j64-emscripten-memory64/LLGo 1428426 B +1554 B / +0.1% (worse) 97789 B 0 B / +0.0%
reflectcall/w32-goos-wasip1/Go 2205707 B 0 B / +0.0% 0 B 0 B / 0.0%
reflectcall/w32-goos-wasip1/LLGo 1548815 B +1994 B / +0.1% (worse) 0 B 0 B / 0.0%
reflectcall/w32-wasi/LLGo 1474699 B +1411 B / +0.1% (worse) 0 B 0 B / 0.0%
w32-goos-wasip1/Go 1909947 B 0 B / +0.0% 0 B 0 B / 0.0%
w32-goos-wasip1/LLGo 147023 B +992 B / +0.7% (worse) 0 B 0 B / 0.0%
w32-wasi/LLGo 147197 B +435 B / +0.3% (worse) 0 B 0 B / 0.0%
LLGo WebAssembly build measurements
Example and profile Build vs base
j32-emscripten 5.910 s +121 ms / +2.1% (worse)
j32-goos-js 5.804 s -146 ms / -2.5% (better)
j64-emscripten-memory64 4.945 s -308.5 us / -0.006238% (better)
reflectcall/w32-wasi 25.261 s +372.2 ms / +1.5% (worse)
w32-goos-wasip1 4.506 s -42.68 ms / -0.9% (better)
w32-wasi 4.509 s +70.71 ms / +1.6% (worse)

Compared with 653957f840b5 measured in the same runner job.

@github-actions

github-actions Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

LLGo baseline benchmarks

f46c1a839521 | workflow run | long-term charts

Program measurements

Platform Workload File size vs base Text size vs base Build vs base Run vs base
Linux cprintf 7160 B 0 B / +0.0% 387 B 0 B / +0.0% 539.523 ms -24.94 ms / -4.4% (better) 1.248 ms -70.03 us / -5.3% (better)
Linux cprintf-lto 6912 B 0 B / +0.0% 368 B 0 B / +0.0% 523.214 ms -56.09 ms / -9.7% (better) 1.274 ms -61.3 us / -4.6% (better)
Linux fmtprintf 1663424 B +104 B / +0.006253% (worse) 498437 B +51 B / +0.01023% (worse) 3.872 s +187.5 ms / +5.1% (worse) 3.200 ms +142.1 us / +4.6% (worse)
Linux fmtprintf-lto 1501168 B +64 B / +0.004264% (worse) 436123 B 0 B / +0.0% 10.634 s -234.3 ms / -2.2% (better) 3.044 ms +111.4 us / +3.8% (worse)
Linux println 69248 B +16 B / +0.02311% (worse) 16805 B 0 B / +0.0% 571.011 ms +7.49 ms / +1.3% (worse) 1.637 ms -76.12 us / -4.4% (better)
Linux println-lto 59976 B +16 B / +0.02668% (worse) 14209 B 0 B / +0.0% 815.326 ms -29.43 ms / -3.5% (better) 1.622 ms -67.36 us / -4.0% (better)
macOS cprintf 68064 B 0 B / +0.0% 4429 B 0 B / +0.0% 1.167 s +197.9 ms / +20.4% (worse) 4.354 ms +1.836 ms / +72.9% (worse)
macOS cprintf-lto 84288 B 0 B / +0.0% 193 B 0 B / +0.0% 1.184 s +185.9 ms / +18.6% (worse) 5.807 ms +1.327 ms / +29.6% (worse)
macOS fmtprintf 1504832 B 0 B / +0.0% 875484 B +164 B / +0.01874% (worse) 5.072 s +388 ms / +8.3% (worse) 7.310 ms +149.8 us / +2.1% (worse)
macOS fmtprintf-lto 1192704 B 0 B / +0.0% 848980 B +28 B / +0.003298% (worse) 9.559 s -2.781 s / -22.5% (better) 6.089 ms -3.395 ms / -35.8% (better)
macOS println 117216 B 0 B / +0.0% 37565 B +16 B / +0.04261% (worse) 1.125 s +239.2 ms / +27.0% (worse) 4.808 ms +883 us / +22.5% (worse)
macOS println-lto 119472 B 0 B / +0.0% 34944 B +16 B / +0.04581% (worse) 1.358 s +83.39 ms / +6.5% (worse) 5.471 ms +893.9 us / +19.5% (worse)
Windows MinGW cprintf 19456 B 0 B / +0.0% 4550 B 0 B / +0.0% 1.072 s -47.37 ms / -4.2% (better) 2.858 ms +42.7 us / +1.5% (worse)
Windows MinGW cprintf-lto 17920 B 0 B / +0.0% 4486 B 0 B / +0.0% 1.095 s -4.317 ms / -0.4% (better) 2.760 ms +25.3 us / +0.9% (worse)
Windows MinGW fmtprintf 1933312 B 0 B / +0.0% 599030 B +80 B / +0.01336% (worse) 3.332 s -51.76 ms / -1.5% (better) 6.102 ms -443 us / -6.8% (better)
Windows MinGW fmtprintf-lto 1957888 B 0 B / +0.0% 547318 B 0 B / +0.0% 8.173 s -181.7 ms / -2.2% (better) 6.319 ms +35.8 us / +0.6% (worse)
Windows MinGW println 75776 B 0 B / +0.0% 25142 B 0 B / +0.0% 1.104 s +15.88 ms / +1.5% (worse) 5.507 ms -344 us / -5.9% (better)
Windows MinGW println-lto 69120 B 0 B / +0.0% 21990 B 0 B / +0.0% 1.276 s -85.27 ms / -6.3% (better) 5.063 ms -1.084 ms / -17.6% (better)
Windows MinGW 386 cprintf 43520 B 0 B / +0.0% 5326 B 0 B / +0.0% 1.286 s -10.37 ms / -0.8% (better) 5.153 ms -492.3 us / -8.7% (better)
Windows MinGW 386 cprintf-lto 20992 B 0 B / +0.0% 5094 B 0 B / +0.0% 1.279 s -35.61 ms / -2.7% (better) 5.082 ms -353.1 us / -6.5% (better)
Windows MinGW 386 fmtprintf 1896448 B 0 B / +0.0% 472462 B +32 B / +0.006773% (worse) 4.209 s +117.4 ms / +2.9% (worse) 11.787 ms +1.939 ms / +19.7% (worse)
Windows MinGW 386 fmtprintf-lto 2181120 B 0 B / +0.0% 451274 B 0 B / +0.0% 9.637 s -20.93 ms / -0.2% (better) 9.979 ms -549.7 us / -5.2% (better)
Windows MinGW 386 println 96256 B 0 B / +0.0% 21474 B 0 B / +0.0% 1.292 s -59.55 ms / -4.4% (better) 8.888 ms -302.7 us / -3.3% (better)
Windows MinGW 386 println-lto 74240 B 0 B / +0.0% 19334 B 0 B / +0.0% 1.479 s -74.18 ms / -4.8% (better) 8.447 ms -242 us / -2.8% (better)
Windows MinGW ARM64 cprintf 18944 B 0 B / +0.0% 4408 B 0 B / +0.0% 1.490 s +4.427 ms / +0.3% (worse) 5.997 ms -53.3 us / -0.9% (better)
Windows MinGW ARM64 cprintf-lto 17920 B 0 B / +0.0% 4340 B 0 B / +0.0% 1.496 s -661.4 us / -0.04418% (better) 5.966 ms -64.4 us / -1.1% (better)
Windows MinGW ARM64 fmtprintf 1819648 B 0 B / +0.0% 509716 B +100 B / +0.01962% (worse) 4.152 s -218.4 ms / -5.0% (better) 12.324 ms -1.359 ms / -9.9% (better)
Windows MinGW ARM64 fmtprintf-lto 1880064 B 0 B / +0.0% 476176 B 0 B / +0.0% 9.599 s -33.66 ms / -0.3% (better) 12.287 ms -859.2 us / -6.5% (better)
Windows MinGW ARM64 println 72192 B 0 B / +0.0% 23884 B 0 B / +0.0% 1.476 s -4.318 ms / -0.3% (better) 10.576 ms +276 us / +2.7% (worse)
Windows MinGW ARM64 println-lto 68608 B 0 B / +0.0% 21232 B 0 B / +0.0% 1.708 s +19.67 ms / +1.2% (worse) 11.053 ms +703.3 us / +6.8% (worse)
Windows MSVC cprintf 120320 B 0 B / +0.0% 65798 B 0 B / +0.0% 857.513 ms +532.5 us / +0.1% (worse) 4.537 ms +1.309 ms / +40.5% (worse)
Windows MSVC cprintf-lto 119808 B 0 B / +0.0% 65734 B 0 B / +0.0% 854.850 ms -14.23 ms / -1.6% (better) 3.324 ms +171.4 us / +5.4% (worse)
Windows MSVC fmtprintf 1643520 B 0 B / +0.0% 694582 B +80 B / +0.01152% (worse) 2.976 s -54.73 ms / -1.8% (better) 11.236 ms +2.62 ms / +30.4% (worse)
Windows MSVC fmtprintf-lto 1635328 B 0 B / +0.0% 647078 B 0 B / +0.0% 7.395 s +27.81 ms / +0.4% (worse) 8.767 ms +488.5 us / +5.9% (worse)
Windows MSVC println 194560 B 0 B / +0.0% 120822 B 0 B / +0.0% 856.135 ms +1.068 ms / +0.1% (worse) 6.516 ms +65.6 us / +1.0% (worse)
Windows MSVC println-lto 192512 B 0 B / +0.0% 118358 B 0 B / +0.0% 1.015 s -20.84 ms / -2.0% (better) 6.604 ms -78.5 us / -1.2% (better)
Windows MSVC 386 cprintf 9728 B 0 B / +0.0% 3931 B 0 B / +0.0% 1.081 s +17.34 ms / +1.6% (worse) 5.777 ms -563.5 us / -8.9% (better)
Windows MSVC 386 cprintf-lto 9216 B 0 B / +0.0% 3853 B 0 B / +0.0% 1.099 s -4.279 ms / -0.4% (better) 6.198 ms +202.1 us / +3.4% (worse)
Windows MSVC 386 fmtprintf 1204736 B 0 B / +0.0% 455852 B +32 B / +0.00702% (worse) 3.805 s +89.92 ms / +2.4% (worse) 11.324 ms -657.3 us / -5.5% (better)
Windows MSVC 386 fmtprintf-lto 1241600 B 0 B / +0.0% 427211 B 0 B / +0.0% 8.746 s +86.94 ms / +1.0% (worse) 12.153 ms +268.8 us / +2.3% (worse)
Windows MSVC 386 println 36352 B 0 B / +0.0% 20324 B 0 B / +0.0% 1.108 s +10.57 ms / +1.0% (worse) 10.322 ms +795.5 us / +8.3% (worse)
Windows MSVC 386 println-lto 35840 B 0 B / +0.0% 18565 B 0 B / +0.0% 1.279 s -18.3 ms / -1.4% (better) 10.219 ms +357.6 us / +3.6% (worse)
Windows MSVC ARM64 cprintf 11776 B 0 B / +0.0% 4192 B 0 B / +0.0% 1.215 s +18.58 ms / +1.6% (worse) 6.390 ms +22.9 us / +0.4% (worse)
Windows MSVC ARM64 cprintf-lto 10752 B 0 B / +0.0% 4084 B 0 B / +0.0% 1.192 s -21.09 ms / -1.7% (better) 6.395 ms -547 us / -7.9% (better)
Windows MSVC ARM64 fmtprintf 1387008 B 0 B / +0.0% 509656 B +96 B / +0.01884% (worse) 3.845 s +57.04 ms / +1.5% (worse) 14.855 ms +992.7 us / +7.2% (worse)
Windows MSVC ARM64 fmtprintf-lto 1405440 B 0 B / +0.0% 476836 B 0 B / +0.0% 8.824 s +115.6 ms / +1.3% (worse) 14.917 ms +140.6 us / +1.0% (worse)
Windows MSVC ARM64 println 45056 B 0 B / +0.0% 23908 B 0 B / +0.0% 1.207 s +12.7 ms / +1.1% (worse) 12.009 ms -313.6 us / -2.5% (better)
Windows MSVC ARM64 println-lto 42496 B 0 B / +0.0% 21380 B 0 B / +0.0% 1.408 s +6.106 ms / +0.4% (worse) 12.214 ms +290.8 us / +2.4% (worse)
Core language and compiler benchmarks
Platform Benchmark ns/op vs base
Linux BenchmarkLookupPCRandom 14.450 ns/op -0.05 ns/op / -0.3% (better)
Linux BenchmarkMergeCompilerFlags 205 ns/op +10 ns/op / +5.1% (worse)
Linux BenchmarkMergeLinkerFlags 126.100 ns/op -9.1 ns/op / -6.7% (better)
Linux BenchmarkChannelBuffered 55.570 ns/op -0.21 ns/op / -0.4% (better)
Linux BenchmarkChannelHandoff 13212 ns/op -494 ns/op / -3.6% (better)
Linux BenchmarkDefer 47.700 ns/op -2.54 ns/op / -5.1% (better)
Linux BenchmarkDirectCall 1.601 ns/op +0.001 ns/op / +0.1% (worse)
Linux BenchmarkGlobalRead 1.193 ns/op +0.023 ns/op / +2.0% (worse)
Linux BenchmarkGlobalWrite 7.773 ns/op -0.024 ns/op / -0.3% (better)
Linux BenchmarkGoroutine 25017 ns/op +930 ns/op / +3.9% (worse)
Linux BenchmarkInterfaceCall 5.883 ns/op -0.167 ns/op / -2.8% (better)
Linux BenchmarkRuntimeGetG 2.943 ns/op -0.019 ns/op / -0.6% (better)
macOS BenchmarkLookupPCRandom 13.250 ns/op -3.54 ns/op / -21.1% (better)
macOS BenchmarkMergeCompilerFlags 148.500 ns/op -13.4 ns/op / -8.3% (better)
macOS BenchmarkMergeLinkerFlags 100.300 ns/op -15.5 ns/op / -13.4% (better)
macOS BenchmarkChannelBuffered 28.460 ns/op -8.37 ns/op / -22.7% (better)
macOS BenchmarkChannelHandoff 7952 ns/op -5127 ns/op / -39.2% (better)
macOS BenchmarkDefer 53.880 ns/op -0.68 ns/op / -1.2% (better)
macOS BenchmarkDirectCall 1.107 ns/op -0.349 ns/op / -24.0% (better)
macOS BenchmarkGlobalRead 1.283 ns/op +0.045 ns/op / +3.6% (worse)
macOS BenchmarkGlobalWrite 1.142 ns/op -0.39 ns/op / -25.5% (better)
macOS BenchmarkGoroutine 81705 ns/op -818 ns/op / -1.0% (better)
macOS BenchmarkInterfaceCall 4.718 ns/op -0.602 ns/op / -11.3% (better)
macOS BenchmarkRuntimeGetG 3.602 ns/op -0.061 ns/op / -1.7% (better)
Windows MinGW BenchmarkLookupPCRandom 9.602 ns/op -0.003 ns/op / -0.03123% (better)
Windows MinGW BenchmarkMergeCompilerFlags 383.900 ns/op +14.2 ns/op / +3.8% (worse)
Windows MinGW BenchmarkMergeLinkerFlags 353.300 ns/op +14.7 ns/op / +4.3% (worse)
Windows MinGW BenchmarkChannelBuffered 24.410 ns/op +0.82 ns/op / +3.5% (worse)
Windows MinGW BenchmarkChannelHandoff 1148 ns/op +115 ns/op / +11.1% (worse)
Windows MinGW BenchmarkDefer 43.100 ns/op -0.48 ns/op / -1.1% (better)
Windows MinGW BenchmarkDirectCall 1.357 ns/op 0 ns/op / +0.0%
Windows MinGW BenchmarkGlobalRead 1.367 ns/op +0.009 ns/op / +0.7% (worse)
Windows MinGW BenchmarkGlobalWrite 2.168 ns/op 0 ns/op / +0.0%
Windows MinGW BenchmarkGoroutine 64500 ns/op +192 ns/op / +0.3% (worse)
Windows MinGW BenchmarkInterfaceCall 6.709 ns/op -0.076 ns/op / -1.1% (better)
Windows MinGW BenchmarkRuntimeGetG 1.631 ns/op +0.002 ns/op / +0.1% (worse)
Windows MinGW 386 BenchmarkLookupPCRandom 26.520 ns/op -0.15 ns/op / -0.6% (better)
Windows MinGW 386 BenchmarkMergeCompilerFlags 767.200 ns/op -18.1 ns/op / -2.3% (better)
Windows MinGW 386 BenchmarkMergeLinkerFlags 696.700 ns/op -21.8 ns/op / -3.0% (better)
Windows MinGW 386 BenchmarkChannelBuffered 38.760 ns/op -0.14 ns/op / -0.4% (better)
Windows MinGW 386 BenchmarkChannelHandoff 859.100 ns/op -98.6 ns/op / -10.3% (better)
Windows MinGW 386 BenchmarkDefer 41.640 ns/op -3.48 ns/op / -7.7% (better)
Windows MinGW 386 BenchmarkDirectCall 1.547 ns/op -0.004 ns/op / -0.3% (better)
Windows MinGW 386 BenchmarkGlobalRead 1.548 ns/op -0.003 ns/op / -0.2% (better)
Windows MinGW 386 BenchmarkGlobalWrite 7.771 ns/op -0.022 ns/op / -0.3% (better)
Windows MinGW 386 BenchmarkGoroutine 108322 ns/op +697 ns/op / +0.6% (worse)
Windows MinGW 386 BenchmarkInterfaceCall 8.432 ns/op +0.026 ns/op / +0.3% (worse)
Windows MinGW 386 BenchmarkRuntimeGetG 1.933 ns/op -0.237 ns/op / -10.9% (better)
Windows MinGW ARM64 BenchmarkLookupPCRandom 12.120 ns/op -0.04 ns/op / -0.3% (better)
Windows MinGW ARM64 BenchmarkMergeCompilerFlags 579.600 ns/op +5.9 ns/op / +1.0% (worse)
Windows MinGW ARM64 BenchmarkMergeLinkerFlags 546.600 ns/op +9.8 ns/op / +1.8% (worse)
Windows MinGW ARM64 BenchmarkChannelBuffered 39 ns/op +0.14 ns/op / +0.4% (worse)
Windows MinGW ARM64 BenchmarkChannelHandoff 2700 ns/op +17 ns/op / +0.6% (worse)
Windows MinGW ARM64 BenchmarkDefer 56.950 ns/op +1.64 ns/op / +3.0% (worse)
Windows MinGW ARM64 BenchmarkDirectCall 0.590 ns/op +0.0002 ns/op / +0.03391% (worse)
Windows MinGW ARM64 BenchmarkGlobalRead 0.663 ns/op 0 ns/op / +0.0%
Windows MinGW ARM64 BenchmarkGlobalWrite 0.663 ns/op -0.0001 ns/op / -0.01507% (better)
Windows MinGW ARM64 BenchmarkGoroutine 64359 ns/op +1398 ns/op / +2.2% (worse)
Windows MinGW ARM64 BenchmarkInterfaceCall 4.147 ns/op +0.001 ns/op / +0.02412% (worse)
Windows MinGW ARM64 BenchmarkRuntimeGetG 1.802 ns/op -0.003 ns/op / -0.2% (better)
Windows MSVC BenchmarkLookupPCRandom 9.722 ns/op -0.101 ns/op / -1.0% (better)
Windows MSVC BenchmarkMergeCompilerFlags 483.100 ns/op +0.4 ns/op / +0.1% (worse)
Windows MSVC BenchmarkMergeLinkerFlags 427.600 ns/op -6.2 ns/op / -1.4% (better)
Windows MSVC BenchmarkChannelBuffered 37.120 ns/op -0.08 ns/op / -0.2% (better)
Windows MSVC BenchmarkChannelHandoff 906.800 ns/op -375.2 ns/op / -29.3% (better)
Windows MSVC BenchmarkDefer 45.130 ns/op +1.26 ns/op / +2.9% (worse)
Windows MSVC BenchmarkDirectCall 1.004 ns/op -0.015 ns/op / -1.5% (better)
Windows MSVC BenchmarkGlobalRead 0.938 ns/op +0.0693 ns/op / +8.0% (worse)
Windows MSVC BenchmarkGlobalWrite 7.179 ns/op +0.027 ns/op / +0.4% (worse)
Windows MSVC BenchmarkGoroutine 68473 ns/op -2619 ns/op / -3.7% (better)
Windows MSVC BenchmarkInterfaceCall 5.107 ns/op -0.377 ns/op / -6.9% (better)
Windows MSVC BenchmarkRuntimeGetG 1.396 ns/op +0.011 ns/op / +0.8% (worse)
Windows MSVC 386 BenchmarkLookupPCRandom 26.520 ns/op -0.08 ns/op / -0.3% (better)
Windows MSVC 386 BenchmarkMergeCompilerFlags 752.300 ns/op +40.9 ns/op / +5.7% (worse)
Windows MSVC 386 BenchmarkMergeLinkerFlags 705 ns/op +22.6 ns/op / +3.3% (worse)
Windows MSVC 386 BenchmarkChannelBuffered 39.290 ns/op -7.03 ns/op / -15.2% (better)
Windows MSVC 386 BenchmarkChannelHandoff 854.200 ns/op +6.1 ns/op / +0.7% (worse)
Windows MSVC 386 BenchmarkDefer 44.270 ns/op -3.43 ns/op / -7.2% (better)
Windows MSVC 386 BenchmarkDirectCall 1.547 ns/op -0.003 ns/op / -0.2% (better)
Windows MSVC 386 BenchmarkGlobalRead 1.858 ns/op -0.005 ns/op / -0.3% (better)
Windows MSVC 386 BenchmarkGlobalWrite 7.791 ns/op -0.009 ns/op / -0.1% (better)
Windows MSVC 386 BenchmarkGoroutine 109872 ns/op +72 ns/op / +0.1% (worse)
Windows MSVC 386 BenchmarkInterfaceCall 8.381 ns/op +0.282 ns/op / +3.5% (worse)
Windows MSVC 386 BenchmarkRuntimeGetG 1.926 ns/op -0.554 ns/op / -22.3% (better)
Windows MSVC ARM64 BenchmarkLookupPCRandom 12.130 ns/op +0.02 ns/op / +0.2% (worse)
Windows MSVC ARM64 BenchmarkMergeCompilerFlags 568 ns/op -2.8 ns/op / -0.5% (better)
Windows MSVC ARM64 BenchmarkMergeLinkerFlags 536.500 ns/op +8.9 ns/op / +1.7% (worse)
Windows MSVC ARM64 BenchmarkChannelBuffered 37.730 ns/op -0.35 ns/op / -0.9% (better)
Windows MSVC ARM64 BenchmarkChannelHandoff 2606 ns/op -749 ns/op / -22.3% (better)
Windows MSVC ARM64 BenchmarkDefer 61.320 ns/op -0.73 ns/op / -1.2% (better)
Windows MSVC ARM64 BenchmarkDirectCall 0.590 ns/op -0.0002 ns/op / -0.03391% (better)
Windows MSVC ARM64 BenchmarkGlobalRead 0.664 ns/op +0.0011 ns/op / +0.2% (worse)
Windows MSVC ARM64 BenchmarkGlobalWrite 3.749 ns/op +0.003 ns/op / +0.1% (worse)
Windows MSVC ARM64 BenchmarkGoroutine 56142 ns/op -2799 ns/op / -4.7% (better)
Windows MSVC ARM64 BenchmarkInterfaceCall 4.137 ns/op -0.001 ns/op / -0.02417% (better)
Windows MSVC ARM64 BenchmarkRuntimeGetG 1.770 ns/op +0.001 ns/op / +0.1% (worse)
Timer runtime benchmarks
Platform Operation and runtime ns/op vs base
Linux AfterFuncZeroDelivery/Go 906.700 ns/op +3.9 ns/op / +0.4% (worse)
Linux AfterFuncZeroDelivery/LLGo 37011 ns/op -10677 ns/op / -22.4% (better)
Linux CreateStop/Go 289.500 ns/op -0.6 ns/op / -0.2% (better)
Linux CreateStop/LLGo 1667 ns/op -15 ns/op / -0.9% (better)
Linux RearmStopped/Go 116.200 ns/op +1 ns/op / +0.9% (worse)
Linux RearmStopped/LLGo 1438 ns/op +224 ns/op / +18.5% (worse)
Linux ResetActive/Go 68.850 ns/op +1.26 ns/op / +1.9% (worse)
Linux ResetActive/LLGo 726.400 ns/op -2.9 ns/op / -0.4% (better)
Linux ResetHeap1024/Go 67.380 ns/op +0.21 ns/op / +0.3% (worse)
Linux ResetHeap1024/LLGo 175.700 ns/op -1.6 ns/op / -0.9% (better)
macOS AfterFuncZeroDelivery/Go 483.500 ns/op -192.1 ns/op / -28.4% (better)
macOS AfterFuncZeroDelivery/LLGo 96682 ns/op -41078 ns/op / -29.8% (better)
macOS CreateStop/Go 152 ns/op -31.8 ns/op / -17.3% (better)
macOS CreateStop/LLGo 511.900 ns/op -536.1 ns/op / -51.2% (better)
macOS RearmStopped/Go 61.580 ns/op -14.12 ns/op / -18.7% (better)
macOS RearmStopped/LLGo 402.500 ns/op -198.5 ns/op / -33.0% (better)
macOS ResetActive/Go 44.680 ns/op -10.47 ns/op / -19.0% (better)
macOS ResetActive/LLGo 180.800 ns/op -68.5 ns/op / -27.5% (better)
macOS ResetHeap1024/Go 46.210 ns/op -8.76 ns/op / -15.9% (better)
macOS ResetHeap1024/LLGo 92.270 ns/op -76.03 ns/op / -45.2% (better)
Windows MinGW AfterFuncZeroDelivery/Go 375.700 ns/op -13.4 ns/op / -3.4% (better)
Windows MinGW AfterFuncZeroDelivery/LLGo 112593 ns/op +797 ns/op / +0.7% (worse)
Windows MinGW CreateStop/Go 89.830 ns/op +0.06 ns/op / +0.1% (worse)
Windows MinGW CreateStop/LLGo 345.300 ns/op -0.7 ns/op / -0.2% (better)
Windows MinGW RearmStopped/Go 24.460 ns/op +0.01 ns/op / +0.0409% (worse)
Windows MinGW RearmStopped/LLGo 228.500 ns/op +5.5 ns/op / +2.5% (worse)
Windows MinGW ResetActive/Go 14.770 ns/op 0 ns/op / +0.0%
Windows MinGW ResetActive/LLGo 123.800 ns/op +2 ns/op / +1.6% (worse)
Windows MinGW ResetHeap1024/Go 14.870 ns/op +0.11 ns/op / +0.7% (worse)
Windows MinGW ResetHeap1024/LLGo 109 ns/op +4 ns/op / +3.8% (worse)
Windows MinGW 386 AfterFuncZeroDelivery/Go 949.500 ns/op -1.4 ns/op / -0.1% (better)
Windows MinGW 386 AfterFuncZeroDelivery/LLGo 201059 ns/op +2534 ns/op / +1.3% (worse)
Windows MinGW 386 CreateStop/Go 191.600 ns/op 0 ns/op / +0.0%
Windows MinGW 386 CreateStop/LLGo 532.600 ns/op +18.4 ns/op / +3.6% (worse)
Windows MinGW 386 RearmStopped/Go 63.770 ns/op +0.61 ns/op / +1.0% (worse)
Windows MinGW 386 RearmStopped/LLGo 340.800 ns/op +1.4 ns/op / +0.4% (worse)
Windows MinGW 386 ResetActive/Go 38.970 ns/op -0.19 ns/op / -0.5% (better)
Windows MinGW 386 ResetActive/LLGo 975.600 ns/op -31.4 ns/op / -3.1% (better)
Windows MinGW 386 ResetHeap1024/Go 39.560 ns/op +0.06 ns/op / +0.2% (worse)
Windows MinGW 386 ResetHeap1024/LLGo 186.900 ns/op -1 ns/op / -0.5% (better)
Windows MinGW ARM64 AfterFuncZeroDelivery/Go 665 ns/op -4.1 ns/op / -0.6% (better)
Windows MinGW ARM64 AfterFuncZeroDelivery/LLGo 146389 ns/op -332 ns/op / -0.2% (better)
Windows MinGW ARM64 CreateStop/Go 204.900 ns/op +8.6 ns/op / +4.4% (worse)
Windows MinGW ARM64 CreateStop/LLGo 354.700 ns/op -15.2 ns/op / -4.1% (better)
Windows MinGW ARM64 RearmStopped/Go 70.610 ns/op -0.02 ns/op / -0.02832% (better)
Windows MinGW ARM64 RearmStopped/LLGo 251.600 ns/op -3.4 ns/op / -1.3% (better)
Windows MinGW ARM64 ResetActive/Go 31.050 ns/op +0.09 ns/op / +0.3% (worse)
Windows MinGW ARM64 ResetActive/LLGo 122.100 ns/op +0.7 ns/op / +0.6% (worse)
Windows MinGW ARM64 ResetHeap1024/Go 31.110 ns/op +0.04 ns/op / +0.1% (worse)
Windows MinGW ARM64 ResetHeap1024/LLGo 127.400 ns/op -0.6 ns/op / -0.5% (better)
Windows MSVC AfterFuncZeroDelivery/Go 582.800 ns/op -15.2 ns/op / -2.5% (better)
Windows MSVC AfterFuncZeroDelivery/LLGo 127811 ns/op -1701 ns/op / -1.3% (better)
Windows MSVC CreateStop/Go 158.600 ns/op -3.7 ns/op / -2.3% (better)
Windows MSVC CreateStop/LLGo 563.800 ns/op -234.7 ns/op / -29.4% (better)
Windows MSVC RearmStopped/Go 59.630 ns/op -0.01 ns/op / -0.01677% (better)
Windows MSVC RearmStopped/LLGo 285.500 ns/op -47.9 ns/op / -14.4% (better)
Windows MSVC ResetActive/Go 26.450 ns/op -0.09 ns/op / -0.3% (better)
Windows MSVC ResetActive/LLGo 152.400 ns/op -29.6 ns/op / -16.3% (better)
Windows MSVC ResetHeap1024/Go 26.740 ns/op 0 ns/op / +0.0%
Windows MSVC ResetHeap1024/LLGo 109.800 ns/op -1.2 ns/op / -1.1% (better)
Windows MSVC 386 AfterFuncZeroDelivery/Go 954.300 ns/op +2.5 ns/op / +0.3% (worse)
Windows MSVC 386 AfterFuncZeroDelivery/LLGo 198834 ns/op -1368 ns/op / -0.7% (better)
Windows MSVC 386 CreateStop/Go 197 ns/op +6.2 ns/op / +3.2% (worse)
Windows MSVC 386 CreateStop/LLGo 451 ns/op +4.9 ns/op / +1.1% (worse)
Windows MSVC 386 RearmStopped/Go 63.380 ns/op +0.05 ns/op / +0.1% (worse)
Windows MSVC 386 RearmStopped/LLGo 321.300 ns/op +10.3 ns/op / +3.3% (worse)
Windows MSVC 386 ResetActive/Go 38.970 ns/op -0.02 ns/op / -0.1% (better)
Windows MSVC 386 ResetActive/LLGo 981.100 ns/op +106.8 ns/op / +12.2% (worse)
Windows MSVC 386 ResetHeap1024/Go 39.540 ns/op +0.19 ns/op / +0.5% (worse)
Windows MSVC 386 ResetHeap1024/LLGo 173.200 ns/op +1.5 ns/op / +0.9% (worse)
Windows MSVC ARM64 AfterFuncZeroDelivery/Go 668 ns/op -4.8 ns/op / -0.7% (better)
Windows MSVC ARM64 AfterFuncZeroDelivery/LLGo 143195 ns/op -19 ns/op / -0.01327% (better)
Windows MSVC ARM64 CreateStop/Go 196.500 ns/op -1.1 ns/op / -0.6% (better)
Windows MSVC ARM64 CreateStop/LLGo 393.400 ns/op +11.6 ns/op / +3.0% (worse)
Windows MSVC ARM64 RearmStopped/Go 70.590 ns/op +0.05 ns/op / +0.1% (worse)
Windows MSVC ARM64 RearmStopped/LLGo 272.200 ns/op -1.2 ns/op / -0.4% (better)
Windows MSVC ARM64 ResetActive/Go 30.970 ns/op -0.14 ns/op / -0.5% (better)
Windows MSVC ARM64 ResetActive/LLGo 137.200 ns/op -7.9 ns/op / -5.4% (better)
Windows MSVC ARM64 ResetHeap1024/Go 31.080 ns/op -0.09 ns/op / -0.3% (better)
Windows MSVC ARM64 ResetHeap1024/LLGo 135.300 ns/op -0.5 ns/op / -0.4% (better)

Compared with 653957f840b5 measured in the same runner job.

@cpunion
cpunion force-pushed the codex/wasi-thread-stability branch from fb2a438 to 4f29c5c Compare September 29, 2026 09:28
@cpunion
cpunion force-pushed the codex/wasi-thread-stability branch 2 times, most recently from 4291ef3 to 5b44983 Compare September 29, 2026 11:34
@cpunion
cpunion force-pushed the codex/wasi-thread-stability branch from 5b44983 to f46c1a8 Compare September 30, 2026 04:31
@cpunion
cpunion merged commit abe36f6 into xgo-dev:main Sep 30, 2026
90 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

wasm/wasi threads: deferred Goexit hangs or exits early under WAMR

2 participants