Skip to content

runtime/wasm: add bounded browser workers with realm-safe callbacks - #2661

Merged
cpunion merged 18 commits into
xgo-dev:mainfrom
cpunion:codex/wasm-workers-config
Sep 28, 2026
Merged

cpunion merged 18 commits into
xgo-dev:mainfrom
cpunion:codex/wasm-workers-config

Conversation

@cpunion

@cpunion cpunion commented Sep 23, 2026 •

Copy link
Copy Markdown
Collaborator

With one browser M, goroutines cannot run in parallel even when Emscripten has Web Workers. LLGO_WASM_WORKERS=2 through 16 now creates a fixed pool of physical workers, each with an M/P and run queue, and schedules many Fiber-backed goroutines per worker. The implementation covers cross-worker wakeups, stop-the-world linear GC, TLS/GLS ownership, finalizers, and Node/Chrome execution for J32 and J64. One-worker mode remains available.

syscall/js.FuncOf keeps synchronous results for nested and nonblocking external callbacks; suspended external callbacks resume after the JS event returns. Emscripten emval handles remain in their owning realm. Descendants of a goroutine using syscall/js stay on that worker; runtime/wasmworkers.GoIndependent explicitly distributes work whose closure has no JS value or thread-local C state. Passing js.Value between unrelated workers remains unsupported, so bounded mode stays opt-in pending that compatibility work (#2632).

The Binaryen pin (#2652) and runner recovery (#2660) are merged. This branch was rebased onto current main (f59bac1ea), leaving only the browser-worker changes in its diff. Before the pin update, the acceptance script passed locally with Binaryen llgo-v132.2 and Emscripten 6.0.8: single-worker J32/J64 and W32 regressions, bounded J32/J64 runtime and hardening fixtures, public llgo test, and four Chrome COOP/COEP runs. The browser-only instantiation-failure test invokes the runner in the mode where it installs instantiateWasm.

With llgo-v132.3, a repeated Memory64 callback test exposed a rare invalid WebAssembly scheduler waiter: the external JS callback or callback poll could take funcsMu on a system Fiber with no G, then try to park when the lock was contended. Bounded mode keeps its poll hook installed, so it now skips the single-worker-only active-callback count on those system-Fiber paths. A new waiter guard reports any future attempt to create a parkable waiter without a G. The failing test passed 300 repeated J64 and 100 J32 runs after the fix; the single-worker callback group passed 50 repetitions.

Review follow-up (083562916): llgo run now preserves the guest's exit code for native and emulator runners; a real exit-code-7 run and focused regression tests passed. LLGO_WASM_WORKERS=1 and the GoIndependent closure restriction are documented. The GC-root and timer build tags now partition consistently, and the shared mutex records waiters so an uncontended unlock does not notify and a contended unlock wakes one waiter. Its bounded 1 ms wait remains necessary for stop-the-world polling when a worker is parked. The complete Node/Chrome J32/J64 worker acceptance passed on this head, as did the focused Go tests. Full PR CI is running.

@fennoai fennoai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review: wasm workers support

This is a large, carefully-engineered change. The _single/_workers/_stub build-tag split is disciplined, the STW epoch protocol and the double-checked wait-queue locking are correct, the lost-wakeup guard (snapshot the wake sequence before checking) is right, and the CI supply-chain hardening (in-repo pinned checksums, regex-validated, verified against a locally-computed hash) is a real improvement. Test coverage is solid and the runtime code is unusually well-commented.

Findings below. The one worth resolving before merge is the llgo run exit-code regression (inline). Everything else is minor / latent / perf-tuning.

Additional (not anchored inline):

  • runtime/internal/runtime/scheduler_events_wasm_workers.go (loadWasmEventHooks) — per-iteration global lock. waitWasmWorkerRunq and cooperativeSafepointSlow call loadWasmEventHooks() on every pass, each acquiring the shared wasmSchedulerEventHooks.lock and copying the struct. The three hook pointers are written only at startup, so every worker's idle loop serializes through one mutex on the busiest path. Publishing/reading the pointers via atomics (or copy-on-write) would make the read lock-free.
  • runtime/internal/lib/syscall/js/emval_release_workers.go — finalizer wake storm. Each releaseEmval fans out WakeWasmCallbackPoll → wakeAllWasmWorkers, so a GC cycle finalizing N JS values wakes every worker N times; pollEmvalReleases also scans the entire pending list per worker (O(W·N)). Batching releases (drain-and-wake once) and using a per-owner queue would remove the per-handle, per-worker storm.
  • runtime/internal/runtime/proc_wasm_workers.go (wakeWasmWorker) — unconditional notify. Wake is issued on every enqueue even when the target worker is already running. The worker.wake sequence counter already provides lost-wake protection, so a "parked" flag would let enqueues skip the notify for running workers (the common case).
  • doc/wasm-proposal.md:88 (and CN mirror ~208) — GoIndependent wording. "starts work in the pool only when its closure carries no JS value or thread-local C state" reads as a runtime decision, but SpawnIndependentWasmG always dispatches; "no JS value / no thread-local C state" is a caller obligation. The Go doc comment in go_workers.go phrases this correctly ("The closure must not carry..."); align the proposal.
  • runtime/internal/runtime/tinygogc/gc_wasm_js_workers.go — empty __libc_free. Intentionally-empty (non-moving GC sweeps) but, unlike the rest of this PR, has no comment. A one-line note would prevent a future "fix".

No high- or medium-severity security findings; the env-var, module-name, path, checksum, and command-execution trust boundaries are all validated correctly. The --no-sandbox Chrome runner is acceptable for CI-scoped, self-built fixtures only — do not reuse it for untrusted content.

Comment thread internal/build/run.go Outdated
Comment thread cmd/internal/run/run.go
Comment thread internal/wasmworkers/config.go
Comment thread runtime/internal/wasmsync/mutex.go
Comment thread runtime/internal/gcroot/current_wasm.go Outdated
Comment thread runtime/internal/lib/runtime/time_wasm_workers_llgo.go
@codecov

codecov Bot commented Sep 23, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 95.58824% with 3 lines in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
internal/build/build.go 89.28% 3 Missing ⚠️

📢 Thoughts on this report? Let us know!

@github-actions

github-actions Bot commented Sep 23, 2026 •

Copy link
Copy Markdown

LLGo WebAssembly build benchmarks

3defce7938c1 | workflow run | long-term charts

WebAssembly output sizes
Example, profile and compiler Wasm module vs base Generated JS glue vs base
cprintf/j32-emscripten/LLGo 146973 B +870 B / +0.6% (worse) 74890 B +58 B / +0.1% (worse)
cprintf/j32-goos-js/LLGo 145407 B +928 B / +0.6% (worse) 73165 B +18 B / +0.02461% (worse)
cprintf/j64-emscripten-memory64/LLGo 134516 B +275 B / +0.2% (worse) 78778 B +58 B / +0.1% (worse)
cprintf/w32-goos-wasip1/LLGo 141956 B +723 B / +0.5% (worse) 0 B 0 B / 0.0%
cprintf/w32-wasi/LLGo 141704 B +721 B / +0.5% (worse) 0 B 0 B / 0.0%
fmtprintf/j32-emscripten/LLGo 3203224 B +15290 B / +0.5% (worse) 118556 B +594 B / +0.5% (worse)
fmtprintf/j32-goos-js/Go 2526852 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/j32-goos-js/LLGo 3182652 B +2699 B / +0.1% (worse) 101635 B +18 B / +0.01771% (worse)
fmtprintf/j64-emscripten-memory64/LLGo 2940192 B +14775 B / +0.5% (worse) 125399 B +58 B / +0.04627% (worse)
fmtprintf/w32-goos-wasip1/Go 2500019 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/w32-goos-wasip1/LLGo 2835788 B +2898 B / +0.1% (worse) 0 B 0 B / 0.0%
fmtprintf/w32-wasi/LLGo 2702221 B +2811 B / +0.1% (worse) 0 B 0 B / 0.0%
j32-emscripten/LLGo 146205 B +867 B / +0.6% (worse) 74890 B +58 B / +0.1% (worse)
j32-goos-js/Go 1895533 B 0 B / +0.0% 0 B 0 B / 0.0%
j32-goos-js/LLGo 144876 B +926 B / +0.6% (worse) 73165 B +18 B / +0.02461% (worse)
j64-emscripten-memory64/LLGo 133847 B +273 B / +0.2% (worse) 78778 B +58 B / +0.1% (worse)
reflectcall/j32-emscripten/LLGo 1533961 B +2674 B / +0.2% (worse) 92056 B +58 B / +0.1% (worse)
reflectcall/j32-goos-js/Go 2191221 B 0 B / +0.0% 0 B 0 B / 0.0%
reflectcall/j32-goos-js/LLGo 1536700 B +2694 B / +0.2% (worse) 90331 B +18 B / +0.01993% (worse)
reflectcall/j64-emscripten-memory64/LLGo 1419158 B +1885 B / +0.1% (worse) 97789 B +58 B / +0.1% (worse)
reflectcall/w32-goos-wasip1/Go 2205707 B 0 B / +0.0% 0 B 0 B / 0.0%
reflectcall/w32-goos-wasip1/LLGo 1540668 B +1987 B / +0.1% (worse) 0 B 0 B / 0.0%
reflectcall/w32-wasi/LLGo 1466211 B +1971 B / +0.1% (worse) 0 B 0 B / 0.0%
w32-goos-wasip1/Go 1909947 B 0 B / +0.0% 0 B 0 B / 0.0%
w32-goos-wasip1/LLGo 141170 B +724 B / +0.5% (worse) 0 B 0 B / 0.0%
w32-wasi/LLGo 140989 B +722 B / +0.5% (worse) 0 B 0 B / 0.0%
LLGo WebAssembly build measurements
Example and profile Build vs base
j32-emscripten 5.829 s -278.4 ms / -4.6% (better)
j32-goos-js 6.010 s -185.1 ms / -3.0% (better)
j64-emscripten-memory64 5.046 s -313 ms / -5.8% (better)
reflectcall/w32-wasi 26.395 s -80.39 ms / -0.3% (better)
w32-goos-wasip1 4.850 s +41.99 ms / +0.9% (worse)
w32-wasi 4.574 s -196.7 ms / -4.1% (better)

Compared with f59bac1ea985 measured in the same runner job.

@github-actions

github-actions Bot commented Sep 23, 2026 •

Copy link
Copy Markdown

LLGo baseline benchmarks

3defce7938c1 | workflow run | long-term charts

Program measurements

Platform Workload File size vs base Text size vs base Build vs base Run vs base
Linux cprintf 7160 B 0 B / +0.0% 387 B 0 B / +0.0% 548.712 ms -32.74 ms / -5.6% (better) 1.270 ms -5.206 us / -0.4% (better)
Linux cprintf-lto 6912 B 0 B / +0.0% 368 B 0 B / +0.0% 548.167 ms -29.8 ms / -5.2% (better) 1.277 ms -39.69 us / -3.0% (better)
Linux fmtprintf 1662208 B -8 B / -0.0004813% (better) 498331 B 0 B / +0.0% 3.974 s -69.69 ms / -1.7% (better) 3.071 ms -34.86 us / -1.1% (better)
Linux fmtprintf-lto 1500368 B 0 B / +0.0% 436107 B 0 B / +0.0% 12.201 s +96.75 ms / +0.8% (worse) 3.116 ms +172.6 us / +5.9% (worse)
Linux println 68992 B 0 B / +0.0% 16783 B 0 B / +0.0% 543.065 ms -42.81 ms / -7.3% (better) 1.595 ms -148.4 us / -8.5% (better)
Linux println-lto 59848 B 0 B / +0.0% 14199 B 0 B / +0.0% 804.019 ms -55.24 ms / -6.4% (better) 1.675 ms +20.24 us / +1.2% (worse)
macOS cprintf 68064 B 0 B / +0.0% 4429 B 0 B / +0.0% 836.709 ms +79.13 ms / +10.4% (worse) 6.372 ms +3.686 ms / +137.2% (worse)
macOS cprintf-lto 84288 B 0 B / +0.0% 193 B 0 B / +0.0% 773.903 ms +44.27 ms / +6.1% (worse) 2.736 ms +312 us / +12.9% (worse)
macOS fmtprintf 1504672 B 0 B / +0.0% 874360 B 0 B / +0.0% 3.161 s -380 ms / -10.7% (better) 5.785 ms +473.9 us / +8.9% (worse)
macOS fmtprintf-lto 1192704 B 0 B / +0.0% 848024 B 0 B / +0.0% 8.125 s +277 ms / +3.5% (worse) 4.678 ms +63.71 us / +1.4% (worse)
macOS println 117136 B 0 B / +0.0% 37421 B 0 B / +0.0% 783.824 ms -7.917 ms / -1.0% (better) 3.867 ms +83.17 us / +2.2% (worse)
macOS println-lto 119472 B 0 B / +0.0% 34824 B 0 B / +0.0% 908.858 ms -169 ms / -15.7% (better) 4.206 ms -125.8 us / -2.9% (better)
Windows MinGW cprintf 19456 B 0 B / +0.0% 4550 B 0 B / +0.0% 1.226 s -16.13 ms / -1.3% (better) 3.367 ms -97 us / -2.8% (better)
Windows MinGW cprintf-lto 17920 B 0 B / +0.0% 4486 B 0 B / +0.0% 1.288 s +35.8 ms / +2.9% (worse) 3.576 ms -34.9 us / -1.0% (better)
Windows MinGW fmtprintf 1932800 B 0 B / +0.0% 598918 B 0 B / +0.0% 4.243 s -2.768 ms / -0.1% (better) 8.295 ms -290.6 us / -3.4% (better)
Windows MinGW fmtprintf-lto 1957376 B 0 B / +0.0% 547318 B 0 B / +0.0% 10.089 s +4.74 ms / +0.047% (worse) 7.807 ms -542.2 us / -6.5% (better)
Windows MinGW println 75776 B 0 B / +0.0% 25142 B 0 B / +0.0% 1.253 s -40.41 ms / -3.1% (better) 7.040 ms +256.8 us / +3.8% (worse)
Windows MinGW println-lto 69120 B 0 B / +0.0% 21990 B 0 B / +0.0% 1.480 s +1.559 ms / +0.1% (worse) 6.918 ms -1.516 ms / -18.0% (better)
Windows MinGW 386 cprintf 43520 B 0 B / +0.0% 5326 B 0 B / +0.0% 1.281 s +26.95 ms / +2.1% (worse) 5.249 ms -284.6 us / -5.1% (better)
Windows MinGW 386 cprintf-lto 20992 B 0 B / +0.0% 5094 B 0 B / +0.0% 1.278 s -33.88 ms / -2.6% (better) 5.222 ms -506.4 us / -8.8% (better)
Windows MinGW 386 fmtprintf 1895936 B 0 B / +0.0% 472398 B 0 B / +0.0% 4.415 s +36.97 ms / +0.8% (worse) 10.779 ms -712.3 us / -6.2% (better)
Windows MinGW 386 fmtprintf-lto 2180608 B 0 B / +0.0% 451258 B 0 B / +0.0% 10.289 s +255.7 ms / +2.5% (worse) 11.448 ms +518.5 us / +4.7% (worse)
Windows MinGW 386 println 96256 B 0 B / +0.0% 21458 B 0 B / +0.0% 1.246 s -28.54 ms / -2.2% (better) 8.622 ms -1.183 ms / -12.1% (better)
Windows MinGW 386 println-lto 74240 B 0 B / +0.0% 19314 B 0 B / +0.0% 1.485 s -12.48 ms / -0.8% (better) 8.834 ms -768.7 us / -8.0% (better)
Windows MinGW ARM64 cprintf 18944 B 0 B / +0.0% 4408 B 0 B / +0.0% 1.584 s +11.31 ms / +0.7% (worse) 7.041 ms +98.4 us / +1.4% (worse)
Windows MinGW ARM64 cprintf-lto 17920 B 0 B / +0.0% 4340 B 0 B / +0.0% 1.632 s +47.07 ms / +3.0% (worse) 6.979 ms +378.2 us / +5.7% (worse)
Windows MinGW ARM64 fmtprintf 1819136 B 0 B / +0.0% 509584 B 0 B / +0.0% 4.603 s +29.96 ms / +0.7% (worse) 14.107 ms +326.9 us / +2.4% (worse)
Windows MinGW ARM64 fmtprintf-lto 1879040 B 0 B / +0.0% 476152 B 0 B / +0.0% 10.436 s +278.2 ms / +2.7% (worse) 14.644 ms +733.1 us / +5.3% (worse)
Windows MinGW ARM64 println 72192 B 0 B / +0.0% 23876 B 0 B / +0.0% 1.601 s +36.58 ms / +2.3% (worse) 11.889 ms +579.5 us / +5.1% (worse)
Windows MinGW ARM64 println-lto 68608 B 0 B / +0.0% 21224 B 0 B / +0.0% 1.821 s +45.92 ms / +2.6% (worse) 12.334 ms +1.119 ms / +10.0% (worse)
Windows MSVC cprintf 120320 B 0 B / +0.0% 65798 B 0 B / +0.0% 942.215 ms +159.8 ms / +20.4% (worse) 3.110 ms +252.7 us / +8.8% (worse)
Windows MSVC cprintf-lto 119808 B 0 B / +0.0% 65734 B 0 B / +0.0% 922.841 ms +128.7 ms / +16.2% (worse) 3.075 ms +169.2 us / +5.8% (worse)
Windows MSVC fmtprintf 1643008 B 0 B / +0.0% 694470 B 0 B / +0.0% 3.173 s +23.9 ms / +0.8% (worse) 8.372 ms +522.7 us / +6.7% (worse)
Windows MSVC fmtprintf-lto 1634816 B 0 B / +0.0% 647062 B 0 B / +0.0% 7.469 s +183.8 ms / +2.5% (worse) 7.792 ms -349.8 us / -4.3% (better)
Windows MSVC println 194560 B 0 B / +0.0% 120822 B 0 B / +0.0% 856.386 ms +71.96 ms / +9.2% (worse) 6.351 ms +539.6 us / +9.3% (worse)
Windows MSVC println-lto 192512 B 0 B / +0.0% 118342 B 0 B / +0.0% 1.007 s +26.68 ms / +2.7% (worse) 6.304 ms +111.9 us / +1.8% (worse)
Windows MSVC 386 cprintf 9728 B 0 B / +0.0% 3931 B 0 B / +0.0% 1.183 s +127.3 ms / +12.1% (worse) 7.806 ms +1.342 ms / +20.8% (worse)
Windows MSVC 386 cprintf-lto 9216 B 0 B / +0.0% 3853 B 0 B / +0.0% 1.068 s -171.9 ms / -13.9% (better) 5.871 ms -142.1 us / -2.4% (better)
Windows MSVC 386 fmtprintf 1204224 B 0 B / +0.0% 455788 B 0 B / +0.0% 3.989 s +11.36 ms / +0.3% (worse) 13.416 ms +836.5 us / +6.6% (worse)
Windows MSVC 386 fmtprintf-lto 1240576 B 0 B / +0.0% 427195 B 0 B / +0.0% 9.083 s -47.54 ms / -0.5% (better) 13.028 ms +36.6 us / +0.3% (worse)
Windows MSVC 386 println 36352 B 0 B / +0.0% 20324 B 0 B / +0.0% 1.076 s -10.2 ms / -0.9% (better) 10.162 ms -40.9 us / -0.4% (better)
Windows MSVC 386 println-lto 35840 B 0 B / +0.0% 18549 B 0 B / +0.0% 1.263 s -8.955 ms / -0.7% (better) 10.311 ms -24.8 us / -0.2% (better)
Windows MSVC ARM64 cprintf 11776 B 0 B / +0.0% 4192 B 0 B / +0.0% 1.207 s -62.34 ms / -4.9% (better) 6.931 ms -1.018 ms / -12.8% (better)
Windows MSVC ARM64 cprintf-lto 10752 B 0 B / +0.0% 4084 B 0 B / +0.0% 1.249 s -42.5 ms / -3.3% (better) 6.707 ms -527.2 us / -7.3% (better)
Windows MSVC ARM64 fmtprintf 1386496 B 0 B / +0.0% 509528 B 0 B / +0.0% 3.865 s -123.1 ms / -3.1% (better) 12.787 ms -797.4 us / -5.9% (better)
Windows MSVC ARM64 fmtprintf-lto 1404928 B 0 B / +0.0% 476820 B 0 B / +0.0% 8.760 s -11.31 ms / -0.1% (better) 14.247 ms -40.9 us / -0.3% (better)
Windows MSVC ARM64 println 45056 B 0 B / +0.0% 23908 B 0 B / +0.0% 1.213 s -30.17 ms / -2.4% (better) 12.355 ms -377 us / -3.0% (better)
Windows MSVC ARM64 println-lto 42496 B 0 B / +0.0% 21380 B 0 B / +0.0% 1.369 s -45.11 ms / -3.2% (better) 13.011 ms +99.8 us / +0.8% (worse)
Core language and compiler benchmarks
Platform Benchmark ns/op vs base
Linux BenchmarkLookupPCRandom 14.690 ns/op +0.15 ns/op / +1.0% (worse)
Linux BenchmarkMergeCompilerFlags 197.800 ns/op -2.1 ns/op / -1.1% (better)
Linux BenchmarkMergeLinkerFlags 127.600 ns/op -8.2 ns/op / -6.0% (better)
Linux BenchmarkChannelBuffered 55.200 ns/op -0.75 ns/op / -1.3% (better)
Linux BenchmarkChannelHandoff 13838 ns/op +587 ns/op / +4.4% (worse)
Linux BenchmarkDefer 53.980 ns/op +2.6 ns/op / +5.1% (worse)
Linux BenchmarkDirectCall 1.184 ns/op +0.011 ns/op / +0.9% (worse)
Linux BenchmarkGlobalRead 1.216 ns/op +0.043 ns/op / +3.7% (worse)
Linux BenchmarkGlobalWrite 7.767 ns/op -0.265 ns/op / -3.3% (better)
Linux BenchmarkGoroutine 23654 ns/op +330 ns/op / +1.4% (worse)
Linux BenchmarkInterfaceCall 5.867 ns/op -0.478 ns/op / -7.5% (better)
Linux BenchmarkRuntimeGetG 2.382 ns/op +0.021 ns/op / +0.9% (worse)
macOS BenchmarkLookupPCRandom 13.490 ns/op -0.26 ns/op / -1.9% (better)
macOS BenchmarkMergeCompilerFlags 108 ns/op -11.9 ns/op / -9.9% (better)
macOS BenchmarkMergeLinkerFlags 76.820 ns/op -2.62 ns/op / -3.3% (better)
macOS BenchmarkChannelBuffered 27.070 ns/op -2.98 ns/op / -9.9% (better)
macOS BenchmarkChannelHandoff 8718 ns/op -33 ns/op / -0.4% (better)
macOS BenchmarkDefer 36.710 ns/op -0.74 ns/op / -2.0% (better)
macOS BenchmarkDirectCall 1.098 ns/op -0.153 ns/op / -12.2% (better)
macOS BenchmarkGlobalRead 1.073 ns/op -0.093 ns/op / -8.0% (better)
macOS BenchmarkGlobalWrite 1.032 ns/op -0.263 ns/op / -20.3% (better)
macOS BenchmarkGoroutine 54022 ns/op +18871 ns/op / +53.7% (worse)
macOS BenchmarkInterfaceCall 4.062 ns/op -0.273 ns/op / -6.3% (better)
macOS BenchmarkRuntimeGetG 2.114 ns/op -0.358 ns/op / -14.5% (better)
Windows MinGW BenchmarkLookupPCRandom 12.990 ns/op -0.02 ns/op / -0.2% (better)
Windows MinGW BenchmarkMergeCompilerFlags 610.800 ns/op -0.7 ns/op / -0.1% (better)
Windows MinGW BenchmarkMergeLinkerFlags 540.300 ns/op +1.2 ns/op / +0.2% (worse)
Windows MinGW BenchmarkChannelBuffered 30.750 ns/op -0.23 ns/op / -0.7% (better)
Windows MinGW BenchmarkChannelHandoff 852.300 ns/op -41 ns/op / -4.6% (better)
Windows MinGW BenchmarkDefer 57.490 ns/op +1.04 ns/op / +1.8% (worse)
Windows MinGW BenchmarkDirectCall 1.549 ns/op +0.001 ns/op / +0.1% (worse)
Windows MinGW BenchmarkGlobalRead 1.856 ns/op -0.01 ns/op / -0.5% (better)
Windows MinGW BenchmarkGlobalWrite 2.458 ns/op -0.001 ns/op / -0.04067% (better)
Windows MinGW BenchmarkGoroutine 86684 ns/op -65 ns/op / -0.1% (better)
Windows MinGW BenchmarkInterfaceCall 8.694 ns/op -0.011 ns/op / -0.1% (better)
Windows MinGW BenchmarkRuntimeGetG 1.859 ns/op -0.001 ns/op / -0.1% (better)
Windows MinGW 386 BenchmarkLookupPCRandom 26.610 ns/op +0.03 ns/op / +0.1% (worse)
Windows MinGW 386 BenchmarkMergeCompilerFlags 771.600 ns/op +26.7 ns/op / +3.6% (worse)
Windows MinGW 386 BenchmarkMergeLinkerFlags 696.100 ns/op -11.9 ns/op / -1.7% (better)
Windows MinGW 386 BenchmarkChannelBuffered 41.760 ns/op +2.42 ns/op / +6.2% (worse)
Windows MinGW 386 BenchmarkChannelHandoff 935.600 ns/op -28.9 ns/op / -3.0% (better)
Windows MinGW 386 BenchmarkDefer 44.250 ns/op +0.85 ns/op / +2.0% (worse)
Windows MinGW 386 BenchmarkDirectCall 1.546 ns/op -0.002 ns/op / -0.1% (better)
Windows MinGW 386 BenchmarkGlobalRead 1.551 ns/op +0.002 ns/op / +0.1% (worse)
Windows MinGW 386 BenchmarkGlobalWrite 7.777 ns/op -0.001 ns/op / -0.01286% (better)
Windows MinGW 386 BenchmarkGoroutine 109592 ns/op +980 ns/op / +0.9% (worse)
Windows MinGW 386 BenchmarkInterfaceCall 8.374 ns/op -0.006 ns/op / -0.1% (better)
Windows MinGW 386 BenchmarkRuntimeGetG 2.168 ns/op -0.003 ns/op / -0.1% (better)
Windows MinGW ARM64 BenchmarkLookupPCRandom 11.980 ns/op -0.1 ns/op / -0.8% (better)
Windows MinGW ARM64 BenchmarkMergeCompilerFlags 556.600 ns/op -13.4 ns/op / -2.4% (better)
Windows MinGW ARM64 BenchmarkMergeLinkerFlags 546 ns/op +15.5 ns/op / +2.9% (worse)
Windows MinGW ARM64 BenchmarkChannelBuffered 37.440 ns/op -1.59 ns/op / -4.1% (better)
Windows MinGW ARM64 BenchmarkChannelHandoff 1699 ns/op -164 ns/op / -8.8% (better)
Windows MinGW ARM64 BenchmarkDefer 57.900 ns/op +0.21 ns/op / +0.4% (worse)
Windows MinGW ARM64 BenchmarkDirectCall 0.590 ns/op -0.0002 ns/op / -0.03392% (better)
Windows MinGW ARM64 BenchmarkGlobalRead 0.663 ns/op -0.0002 ns/op / -0.03016% (better)
Windows MinGW ARM64 BenchmarkGlobalWrite 0.663 ns/op -0.0006 ns/op / -0.1% (better)
Windows MinGW ARM64 BenchmarkGoroutine 61355 ns/op -827 ns/op / -1.3% (better)
Windows MinGW ARM64 BenchmarkInterfaceCall 4.234 ns/op +0.003 ns/op / +0.1% (worse)
Windows MinGW ARM64 BenchmarkRuntimeGetG 1.799 ns/op +0.03 ns/op / +1.7% (worse)
Windows MSVC BenchmarkLookupPCRandom 9.071 ns/op +0.648 ns/op / +7.7% (worse)
Windows MSVC BenchmarkMergeCompilerFlags 445 ns/op +54.1 ns/op / +13.8% (worse)
Windows MSVC BenchmarkMergeLinkerFlags 393.400 ns/op +32 ns/op / +8.9% (worse)
Windows MSVC BenchmarkChannelBuffered 29.120 ns/op -1.32 ns/op / -4.3% (better)
Windows MSVC BenchmarkChannelHandoff 3832 ns/op +1517 ns/op / +65.5% (worse)
Windows MSVC BenchmarkDefer 38.590 ns/op -6.88 ns/op / -15.1% (better)
Windows MSVC BenchmarkDirectCall 0.260 ns/op -0.0136 ns/op / -5.0% (better)
Windows MSVC BenchmarkGlobalRead 0.360 ns/op +0.0102 ns/op / +2.9% (worse)
Windows MSVC BenchmarkGlobalWrite 7.185 ns/op +0.237 ns/op / +3.4% (worse)
Windows MSVC BenchmarkGoroutine 64792 ns/op -4809 ns/op / -6.9% (better)
Windows MSVC BenchmarkInterfaceCall 4.335 ns/op +0.016 ns/op / +0.4% (worse)
Windows MSVC BenchmarkRuntimeGetG 0.918 ns/op +0.0054 ns/op / +0.6% (worse)
Windows MSVC 386 BenchmarkLookupPCRandom 69.870 ns/op +0.05 ns/op / +0.1% (worse)
Windows MSVC 386 BenchmarkMergeCompilerFlags 791 ns/op +11 ns/op / +1.4% (worse)
Windows MSVC 386 BenchmarkMergeLinkerFlags 728.900 ns/op -35.2 ns/op / -4.6% (better)
Windows MSVC 386 BenchmarkChannelBuffered 53.300 ns/op -0.91 ns/op / -1.7% (better)
Windows MSVC 386 BenchmarkChannelHandoff 2235 ns/op -1848 ns/op / -45.3% (better)
Windows MSVC 386 BenchmarkDefer 52.550 ns/op +0.58 ns/op / +1.1% (worse)
Windows MSVC 386 BenchmarkDirectCall 1.083 ns/op +0.015 ns/op / +1.4% (worse)
Windows MSVC 386 BenchmarkGlobalRead 1.345 ns/op -0.008 ns/op / -0.6% (better)
Windows MSVC 386 BenchmarkGlobalWrite 17.200 ns/op -0.01 ns/op / -0.1% (better)
Windows MSVC 386 BenchmarkGoroutine 245745 ns/op +2649 ns/op / +1.1% (worse)
Windows MSVC 386 BenchmarkInterfaceCall 5.778 ns/op +0.202 ns/op / +3.6% (worse)
Windows MSVC 386 BenchmarkRuntimeGetG 1.553 ns/op +0.088 ns/op / +6.0% (worse)
Windows MSVC ARM64 BenchmarkLookupPCRandom 12.080 ns/op +0.04 ns/op / +0.3% (worse)
Windows MSVC ARM64 BenchmarkMergeCompilerFlags 572.200 ns/op +8.8 ns/op / +1.6% (worse)
Windows MSVC ARM64 BenchmarkMergeLinkerFlags 524.700 ns/op -4.1 ns/op / -0.8% (better)
Windows MSVC ARM64 BenchmarkChannelBuffered 38.890 ns/op -0.09 ns/op / -0.2% (better)
Windows MSVC ARM64 BenchmarkChannelHandoff 1773 ns/op -174 ns/op / -8.9% (better)
Windows MSVC ARM64 BenchmarkDefer 61.590 ns/op -0.14 ns/op / -0.2% (better)
Windows MSVC ARM64 BenchmarkDirectCall 0.590 ns/op +0.0003 ns/op / +0.1% (worse)
Windows MSVC ARM64 BenchmarkGlobalRead 0.664 ns/op +0.0002 ns/op / +0.03015% (worse)
Windows MSVC ARM64 BenchmarkGlobalWrite 3.758 ns/op +0.006 ns/op / +0.2% (worse)
Windows MSVC ARM64 BenchmarkGoroutine 53574 ns/op -1475 ns/op / -2.7% (better)
Windows MSVC ARM64 BenchmarkInterfaceCall 4.239 ns/op +0.003 ns/op / +0.1% (worse)
Windows MSVC ARM64 BenchmarkRuntimeGetG 1.783 ns/op +0.014 ns/op / +0.8% (worse)
Timer runtime benchmarks
Platform Operation and runtime ns/op vs base
Linux AfterFuncZeroDelivery/Go 903.600 ns/op -23.9 ns/op / -2.6% (better)
Linux AfterFuncZeroDelivery/LLGo 37074 ns/op -3873 ns/op / -9.5% (better)
Linux CreateStop/Go 290.600 ns/op -1.8 ns/op / -0.6% (better)
Linux CreateStop/LLGo 1897 ns/op +61 ns/op / +3.3% (worse)
Linux RearmStopped/Go 116.700 ns/op +1 ns/op / +0.9% (worse)
Linux RearmStopped/LLGo 1349 ns/op +50 ns/op / +3.8% (worse)
Linux ResetActive/Go 68.670 ns/op +1.09 ns/op / +1.6% (worse)
Linux ResetActive/LLGo 801.800 ns/op +0.6 ns/op / +0.1% (worse)
Linux ResetHeap1024/Go 67.160 ns/op -0.23 ns/op / -0.3% (better)
Linux ResetHeap1024/LLGo 175.200 ns/op +0.8 ns/op / +0.5% (worse)
macOS AfterFuncZeroDelivery/Go 412.300 ns/op -81.2 ns/op / -16.5% (better)
macOS AfterFuncZeroDelivery/LLGo 66771 ns/op -6305 ns/op / -8.6% (better)
macOS CreateStop/Go 130.200 ns/op -24 ns/op / -15.6% (better)
macOS CreateStop/LLGo 480.300 ns/op -53 ns/op / -9.9% (better)
macOS RearmStopped/Go 58.430 ns/op -7.91 ns/op / -11.9% (better)
macOS RearmStopped/LLGo 381 ns/op -40.2 ns/op / -9.5% (better)
macOS ResetActive/Go 43.900 ns/op -3.71 ns/op / -7.8% (better)
macOS ResetActive/LLGo 176.100 ns/op -24.2 ns/op / -12.1% (better)
macOS ResetHeap1024/Go 41.370 ns/op -5.8 ns/op / -12.3% (better)
macOS ResetHeap1024/LLGo 84.840 ns/op -6.99 ns/op / -7.6% (better)
Windows MinGW AfterFuncZeroDelivery/Go 551.200 ns/op -29.4 ns/op / -5.1% (better)
Windows MinGW AfterFuncZeroDelivery/LLGo 179303 ns/op +4617 ns/op / +2.6% (worse)
Windows MinGW CreateStop/Go 117.800 ns/op +2.2 ns/op / +1.9% (worse)
Windows MinGW CreateStop/LLGo 446.600 ns/op +13.7 ns/op / +3.2% (worse)
Windows MinGW RearmStopped/Go 31.310 ns/op 0 ns/op / +0.0%
Windows MinGW RearmStopped/LLGo 263.600 ns/op -24.3 ns/op / -8.4% (better)
Windows MinGW ResetActive/Go 20.110 ns/op -0.15 ns/op / -0.7% (better)
Windows MinGW ResetActive/LLGo 156 ns/op -8.4 ns/op / -5.1% (better)
Windows MinGW ResetHeap1024/Go 20.520 ns/op +0.16 ns/op / +0.8% (worse)
Windows MinGW ResetHeap1024/LLGo 127.900 ns/op +2.6 ns/op / +2.1% (worse)
Windows MinGW 386 AfterFuncZeroDelivery/Go 960.800 ns/op +8.4 ns/op / +0.9% (worse)
Windows MinGW 386 AfterFuncZeroDelivery/LLGo 202881 ns/op +1803 ns/op / +0.9% (worse)
Windows MinGW 386 CreateStop/Go 193.300 ns/op +2.9 ns/op / +1.5% (worse)
Windows MinGW 386 CreateStop/LLGo 490.200 ns/op -8.5 ns/op / -1.7% (better)
Windows MinGW 386 RearmStopped/Go 63.680 ns/op +0.5 ns/op / +0.8% (worse)
Windows MinGW 386 RearmStopped/LLGo 338.900 ns/op +1 ns/op / +0.3% (worse)
Windows MinGW 386 ResetActive/Go 39.100 ns/op +0.05 ns/op / +0.1% (worse)
Windows MinGW 386 ResetActive/LLGo 995.100 ns/op +21.2 ns/op / +2.2% (worse)
Windows MinGW 386 ResetHeap1024/Go 39.470 ns/op +0.11 ns/op / +0.3% (worse)
Windows MinGW 386 ResetHeap1024/LLGo 188.300 ns/op -1.1 ns/op / -0.6% (better)
Windows MinGW ARM64 AfterFuncZeroDelivery/Go 661 ns/op -4.9 ns/op / -0.7% (better)
Windows MinGW ARM64 AfterFuncZeroDelivery/LLGo 151454 ns/op -12075 ns/op / -7.4% (better)
Windows MinGW ARM64 CreateStop/Go 194.500 ns/op -2 ns/op / -1.0% (better)
Windows MinGW ARM64 CreateStop/LLGo 399.300 ns/op -14.9 ns/op / -3.6% (better)
Windows MinGW ARM64 RearmStopped/Go 70.610 ns/op +0.04 ns/op / +0.1% (worse)
Windows MinGW ARM64 RearmStopped/LLGo 264.400 ns/op +3.3 ns/op / +1.3% (worse)
Windows MinGW ARM64 ResetActive/Go 31.160 ns/op +0.03 ns/op / +0.1% (worse)
Windows MinGW ARM64 ResetActive/LLGo 137.600 ns/op -0.8 ns/op / -0.6% (better)
Windows MinGW ARM64 ResetHeap1024/Go 31.120 ns/op +0.07 ns/op / +0.2% (worse)
Windows MinGW ARM64 ResetHeap1024/LLGo 127.200 ns/op +0.1 ns/op / +0.1% (worse)
Windows MSVC AfterFuncZeroDelivery/Go 474.800 ns/op -0.8 ns/op / -0.2% (better)
Windows MSVC AfterFuncZeroDelivery/LLGo 134024 ns/op -5624 ns/op / -4.0% (better)
Windows MSVC CreateStop/Go 134 ns/op -10.9 ns/op / -7.5% (better)
Windows MSVC CreateStop/LLGo 638.100 ns/op +16.4 ns/op / +2.6% (worse)
Windows MSVC RearmStopped/Go 55.970 ns/op +4.92 ns/op / +9.6% (worse)
Windows MSVC RearmStopped/LLGo 241.700 ns/op +2.3 ns/op / +1.0% (worse)
Windows MSVC ResetActive/Go 22.550 ns/op -1.55 ns/op / -6.4% (better)
Windows MSVC ResetActive/LLGo 165.500 ns/op -54.6 ns/op / -24.8% (better)
Windows MSVC ResetHeap1024/Go 22.500 ns/op -0.07 ns/op / -0.3% (better)
Windows MSVC ResetHeap1024/LLGo 94.440 ns/op -0.28 ns/op / -0.3% (better)
Windows MSVC 386 AfterFuncZeroDelivery/Go 1005 ns/op -3 ns/op / -0.3% (better)
Windows MSVC 386 AfterFuncZeroDelivery/LLGo 198802 ns/op -1277 ns/op / -0.6% (better)
Windows MSVC 386 CreateStop/Go 265.300 ns/op +8.4 ns/op / +3.3% (worse)
Windows MSVC 386 CreateStop/LLGo 723.600 ns/op -256.9 ns/op / -26.2% (better)
Windows MSVC 386 RearmStopped/Go 93.320 ns/op -0.27 ns/op / -0.3% (better)
Windows MSVC 386 RearmStopped/LLGo 417 ns/op +21.3 ns/op / +5.4% (worse)
Windows MSVC 386 ResetActive/Go 45.150 ns/op -0.1 ns/op / -0.2% (better)
Windows MSVC 386 ResetActive/LLGo 319.600 ns/op -60.1 ns/op / -15.8% (better)
Windows MSVC 386 ResetHeap1024/Go 45.510 ns/op +0.05 ns/op / +0.1% (worse)
Windows MSVC 386 ResetHeap1024/LLGo 169.500 ns/op +2.6 ns/op / +1.6% (worse)
Windows MSVC ARM64 AfterFuncZeroDelivery/Go 655.400 ns/op -20.8 ns/op / -3.1% (better)
Windows MSVC ARM64 AfterFuncZeroDelivery/LLGo 155658 ns/op +9353 ns/op / +6.4% (worse)
Windows MSVC ARM64 CreateStop/Go 192.700 ns/op -7.3 ns/op / -3.7% (better)
Windows MSVC ARM64 CreateStop/LLGo 470.900 ns/op +6.6 ns/op / +1.4% (worse)
Windows MSVC ARM64 RearmStopped/Go 70.660 ns/op +0.1 ns/op / +0.1% (worse)
Windows MSVC ARM64 RearmStopped/LLGo 280.200 ns/op +3.9 ns/op / +1.4% (worse)
Windows MSVC ARM64 ResetActive/Go 31.050 ns/op -0.06 ns/op / -0.2% (better)
Windows MSVC ARM64 ResetActive/LLGo 148.200 ns/op -6.8 ns/op / -4.4% (better)
Windows MSVC ARM64 ResetHeap1024/Go 31.090 ns/op +0.02 ns/op / +0.1% (worse)
Windows MSVC ARM64 ResetHeap1024/LLGo 137.100 ns/op +0.7 ns/op / +0.5% (worse)

Compared with f59bac1ea985 measured in the same runner job.

@cpunion
cpunion force-pushed the codex/wasm-workers-config branch from 48d850e to 0f7cb03 Compare September 27, 2026 15:06
@cpunion

cpunion commented Sep 27, 2026

Copy link
Copy Markdown
Collaborator Author

Thanks for the detailed review. Commit 083562916 restores the guest exit code for both native and emulator llgo run paths (including a real exit-7 run), documents LLGO_WASM_WORKERS=1 and the GoIndependent caller obligation, fixes both build-tag gaps, and explains the intentionally empty __libc_free. The mutex now tracks waiters and wakes one only when contended. Its 1 ms bounded wait remains necessary for stop-the-world polling while a worker is parked. The timeout-help issue was fixed in merged #2660. The complete Node/Chrome J32/J64 worker acceptance and focused Go tests pass on this head.

The event-hook read lock, emval-release fan-out, and worker wake-on-enqueue are valid performance follow-ups. I am leaving those out of this functional PR pending contention measurements, rather than changing three scheduler paths without evidence of the best batching/parking strategy.

@visualfc visualfc left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Follow-up review on the current head. The earlier blocking item (guest exit code on llgo run) is fixed, and the STW / waiter / realm-affinity design still looks sound. Two functional gaps below; neither needs to block this opt-in merge.

The event-hook read lock, emval-release wake storm, and unconditional worker notify remain valid performance follow-ups.

Comment thread runtime/internal/runtime/proc_wasm_workers.go Outdated
Comment thread runtime/_patch/syscall/fs_js_wasm_workers.go
@cpunion
cpunion merged commit 8b8ffc1 into xgo-dev:main Sep 28, 2026
88 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants