Skip to content

wasm: make reflection entries callable across workers - #2743

Open
cpunion wants to merge 13 commits into
xgo-dev:mainfrom
cpunion:codex/wasm-reflect-worker-tables-20261006
Open

cpunion wants to merge 13 commits into
xgo-dev:mainfrom
cpunion:codex/wasm-reflect-worker-tables-20261006

Conversation

@cpunion

@cpunion cpunion commented Oct 6, 2026 •

Copy link
Copy Markdown
Collaborator

Builds on #2738. reflect.MakeFunc goroutines trapped with “table index is out of bounds” because libffi installed a dynamic entry only in the creating worker's table. Reuse compiler-generated Wasm reflection entries for JavaScript providers so every instance has the same callable entry. Native reflection and C libffi retain their backends; remove obsolete Go-reflection JS trampolines/Asyncify exclusions. Preserve reachability gating, signature deduplication and cache fingerprints.

Validation covers old-backend failure/new-backend pass, reflect/defer/recover, real profile/group/cache tests, and Node/Chrome Memory32/64 with 2/4 workers. Worker CI runs the complete test/go package with its original 60-second limit. The standalone reflection fixture is restricted to Wasm workers. The parent fixes allocator progress, bounded GC scanning and rooted filesystem buffers; full Memory32 passes locally around 20–21 seconds. Rebased onto the parent follow-ups and latest main be669acd8; current hosted CI and new-source compatibility qualification remain pending.

Size tradeoff on the minimal workload: Wasm grows about 62 KB and JS shrinks about 13 KB. Large test/go grows from about 15.75 MB to 18.20 MB; bridge/metadata size optimization remains tracked in #2679.

@cpunion

cpunion commented Oct 6, 2026

Copy link
Copy Markdown
Collaborator Author

@fennoai please review the reflection follow-up range 87995c1..958f455, especially worker-local Wasm table correctness, typed bridge reachability/cache consistency and native backend isolation. Full std qualification reproduced the old backend trap on both Memory32/64. Existing reflection/defer/recover tests and compiler/cache tests pass with the fix; the new Node/Chrome worker fixture is running locally. The PR records the measured module size cost rather than claiming this backend change reduces size.

@fennoai fennoai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review summary

This PR reworks WASM Emscripten worker mode so syscall/js values are shared across workers (all remote JS ops forward to worker zero via a per-caller executor G), and switches Emscripten reflection from libffi to compiler-generated typed bridges so MakeFunc values survive crossing workers. The engineering is careful and well-commented: build-tag partitioning is mutually exclusive and exhaustive, the executor reuse / //go:noinline pattern is deliberate, the runqueue Remove/RemoveAfter/Front additions and the finalizer-mutex refactor are correct, and the bilingual proposal/doc updates match the code with no stale jsRealm/MarkCurrentJSRealm/libffi-reflection references remaining.

Findings below. The scheduler items (reentrancy + O(n) child scan) are the ones worth a close look; the rest are minor.

Body-level finding (location not inline-commentable):

Possible reentrant goschedBackend while a G is mid status-transition. goschedBackend (runtime/internal/runtime/proc_wasm_workers.go:399-408) transitions gp to _Grunnable via casgstatus(gp, _Grunning, _Grunnable) (line 402) before acquiring the worker lock in enqueueWasmG (line 403 → worker.lock.Lock(CooperativeSafepoint) at 449). On a contended lock, wasmsync.Mutex.Lock invokes the yield (CooperativeSafepoint). If the safepoint budget is exhausted and no GC stop is pending, cooperativeSafepointSlow (runtime/internal/runtime/safepoint_wasm_workers.go:63-65, not in this diff) calls goschedBackend() again when the runq is non-empty, re-running casgstatus(gp, _Grunning, _Grunnable) on a G that is already _Grunnable — a failed transition that would fatal. goready (line 431-433) has the same shape. This is narrow (needs same-worker lock contention + budget expiry at that exact poll + non-empty runq) and may be unreachable for reasons I can't fully confirm without running the scheduler, but since this PR makes cross-worker enqueues (and thus lock contention) much more frequent, it's worth confirming the queue-lock yield cannot reenter goschedBackend while a caller holds gp in a transient status.

Comment thread runtime/internal/runtime/proc_wasm_workers.go
Comment thread runtime/internal/lib/syscall/js/worker_calls.go
Comment thread runtime/internal/runtime/js_worker_call_wasm_workers.go
Comment thread runtime/internal/runtime/js_worker_call_wasm_workers.go Outdated
@fennoai

fennoai Bot commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

Reflection follow-up review — 87995c126..958f4552a

Focused on the three areas you called out. The backend swap (libffi JS trampolines → compiler-generated typed bridges for all GOOS=js providers) is coherent and the test matrix changes correctly pin the new behavior. One concrete cleanup finding in the native-backend-isolation area, plus a cosmetic consistency nit.

Finding — stale Emscripten reflection linker flags (native backend isolation)

internal/crosscompile/crosscompile.go still passes libffi-reflection glue that this commit removed for the js profile:

  • emscriptenAsyncifyRemove (line 129, applied at line 886):
    -sASYNCIFY_REMOVE=["*llgo_reflect_bind0_js*","*llgo_reflect_bind1_js*","*llgo_reflect_bindn_js*"]
    Those three C trampolines were defined only in the now-deleted runtime/internal/lib/reflect/_wrap/makefunc_wasm_js.c. They no longer exist in any js build, so these ASYNCIFY_REMOVE patterns now match nothing.
  • emscriptenAsyncifyImports (line ~124): still lists ffi_call_js, described in its comment as the import needed so "Asyncify cannot unwind a sleeping Go function invoked by reflect.Value.Call / MakeFunc." That path is now the typed bridge (callWasmBridge), which doesn't route reflection through ffi_call_js (no definition/use of that symbol remains in-tree).

Impact is low — unmatched ASYNCIFY_REMOVE wildcards are effectively a no-op and a spurious async import just keeps an unused symbol unwindable — so this won't break the build. But both flags, and their adjacent comments (lines 122–128), now describe a reflection mechanism that no longer applies to this profile, which is misleading for the next person touching the Emscripten link step. Worth removing the reflection-trampoline entries (and trimming ffi_call_js from imports if it's only reflection-related) or re-scoping the comments to the remaining C-integration libffi use. Note appendEmscriptenLibffiSearchPath and the C libffi path are still legitimately needed, so libffi itself should stay.

Nit — predicate form differs across the three gates (cache consistency)

The reflect-bridge gate is now spelled three ways: target.GOOS == "js" in initial_groups.go:46 and wasm_reflect.go:43 (raw field) vs p.effectiveGOOS() == "js" in ssa/target.go:62 (host-GOOS fallback when empty). For wasm builds GOOS is always set explicitly, so they classify identically today and the group cache key stays consistent with what's emitted — but aligning all three on the same helper would remove a latent way for the build-group key and the actual bridge emission to diverge.

Verified consistent (no action)

  • Group dedup key (initial_groups.go) agrees with configureWasmReflectBridges; reachability gating via usesWasmReflectBridges() is preserved, so dead-reflection std helpers still emit nothing per build.
  • Build-tag pairs for functype_*, call_bridge_*, makefunc_ffi.go are exact complements (wasip1 || js vs !wasip1 && !js); no js-profile file still references the deleted bind0JS/bind1JS/bindnJS/invokeJSMakeFunc symbols.
  • Worker-local-table rationale is sound, and the new test/wasm/reflect-workers/main.go fixture has no build constraint, so it exercises the typed-bridge path under single-worker js and WASI too, not only worker mode — good coverage for the widened gate.
  • Test matrices (wasm_reflect_test.go, reflect_bridge_test.go) correctly flip GoJS J32 to typed bridges and add active Emscripten J32/J64 cases.

Thanks for the context on the reproduced old-backend trap and the module-size-cost recording — that matches what the proposal text now claims (measured cost, not a size-reduction claim).

@cpunion
cpunion force-pushed the codex/wasm-reflect-worker-tables-20261006 branch from 958f455 to f54fe33 Compare October 6, 2026 05:33
@cpunion

cpunion commented Oct 6, 2026

Copy link
Copy Markdown
Collaborator Author

The body-level runqueue finding was confirmed with a controlled contended-lock/budget-exhaustion probe: the preemptive callback fails with invalid goroutine status transition; the in-place GC callback passes. Parent #2738 now uses the existing in-place GC wait at all three queue-lock sites (cc47f7a), with a regression exercising actual Gosched under contention both without and with concurrent GC. This PR is rebased onto that fix. The reflection changes also remove the obsolete Asyncify trampoline exclusions; C libffi imports remain supported.

@cpunion

cpunion commented Oct 6, 2026

Copy link
Copy Markdown
Collaborator Author

@fennoai please review only the reflection follow-up range cc47f7a..f54fe33; cc47f7a is the parent #2738 and contains the scheduler fixes discussed above. The follow-up replaces worker-local dynamic reflection entries with shared static ones, keeps reachability/cache gating, and removes obsolete reflection-only Asyncify exclusions. Full local Node/Chrome 2/4-worker Memory32/64 acceptance passes; fresh std qualification is pinned to the reflection fix at 958f455. Native/C libffi backends remain supported. The outstanding scan-complexity discussion belongs to the parent callback scheduler and is recorded as performance qualification rather than changed by this reflection patch.

@fennoai

fennoai Bot commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

Reflection follow-up review — cc47f7a1f..f54fe33ba

Reviewed the two-commit range (b67824efe compiled reflection entries, reparented onto #2738; f54fe33ba Asyncify-exclusion cleanup). The stale-flags finding from the previous round is fully resolved, and no new issues.

Resolved — obsolete reflection-only Asyncify exclusions removed ✓

f54fe33ba addresses exactly what I flagged last round:

  • Drops the emscriptenAsyncifyRemove constant and its LDFLAG — the *llgo_reflect_bind0_js*/bind1/bindn patterns pointed at trampolines that lived only in the deleted _wrap/makefunc_wasm_js.c, so they matched nothing.
  • Removes the now-dead crosscompile_test.go assertion for that flag.
  • Re-scopes the ffi_call_js comment to "the C libffi API. Go reflection uses static typed entries" — correctly separating the retained C-integration use from the removed reflection use.

Verified:

  • No dangling references anywhere to the removed trampoline symbols or ASYNCIFY_REMOVE (grep across .go/.c is clean).
  • ffi_call_js correctly retained in ASYNCIFY_IMPORTS — libffi's ffi_call is still used by the C path (runtime/internal/clite/ffi/_wrap/libffi.c:280), so the native/C backend stays unwindable. Matches your note that native/C libffi backends remain supported.
  • Changed Go files are gofmt-clean, including the realigned const block after removing the longest-named constant.

Unchanged — reflection-entries commit matches prior-reviewed content ✓

b67824efe is the same 16-file change I reviewed earlier (shared static typed bridges for all GOOS=js providers, reachability/cache gating preserved, build-tag complements exact, reflect-workers fixture unconstrained so it also exercises single-worker js + WASI). No new concerns.

Nit (carryover, non-blocking)

The reflect-bridge gate is still spelled target.GOOS == "js" in initial_groups.go:46 + wasm_reflect.go:43 but p.effectiveGOOS() == "js" in ssa/target.go:62. Identical in practice (wasm always sets GOOS explicitly), so the group cache key stays consistent with emitted bridges — purely a style-alignment opportunity, not a defect.

Noted and agreed that the callback-scheduler scan-complexity item belongs to the parent (#2738) and is tracked as performance qualification, not something this reflection patch changes.

@codecov

codecov Bot commented Oct 6, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@github-actions

github-actions Bot commented Oct 6, 2026 •

Copy link
Copy Markdown

LLGo WebAssembly build benchmarks

218907942ded | workflow run | long-term charts

WebAssembly output sizes
Example, profile and compiler Wasm module vs base Generated JS glue vs base
cprintf/j32-emscripten/LLGo 149901 B +661 B / +0.4% (worse) 88541 B +384 B / +0.4% (worse)
cprintf/j32-goos-js/LLGo 148754 B +243 B / +0.2% (worse) 72985 B 0 B / +0.0%
cprintf/j64-emscripten-memory64/LLGo 137240 B +533 B / +0.4% (worse) 91712 B +384 B / +0.4% (worse)
cprintf/w32-goos-wasip1/LLGo 154964 B +668 B / +0.4% (worse) 0 B 0 B / 0.0%
cprintf/w32-wasi/LLGo 155667 B +688 B / +0.4% (worse) 0 B 0 B / 0.0%
fmtprintf/j32-emscripten/LLGo 3042191 B -14808 B / -0.5% (better) 120701 B -9672 B / -7.4% (better)
fmtprintf/j32-goos-js/Go 2526852 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/j32-goos-js/LLGo 3134304 B -18675 B / -0.6% (better) 90715 B -10058 B / -10.0% (better)
fmtprintf/j64-emscripten-memory64/LLGo 2806585 B -12507 B / -0.4% (better) 125242 B -10611 B / -7.8% (better)
fmtprintf/w32-goos-wasip1/Go 2500019 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/w32-goos-wasip1/LLGo 2356403 B +955 B / +0.04054% (worse) 0 B 0 B / 0.0%
fmtprintf/w32-wasi/LLGo 2353001 B +975 B / +0.04145% (worse) 0 B 0 B / 0.0%
j32-emscripten/LLGo 149245 B +663 B / +0.4% (worse) 88541 B +384 B / +0.4% (worse)
j32-goos-js/Go 1895533 B 0 B / +0.0% 0 B 0 B / 0.0%
j32-goos-js/LLGo 148247 B +245 B / +0.2% (worse) 72985 B 0 B / +0.0%
j64-emscripten-memory64/LLGo 136577 B +530 B / +0.4% (worse) 91712 B +384 B / +0.4% (worse)
reflectcall/j32-emscripten/LLGo 1546105 B +64982 B / +4.4% (worse) 95329 B -9588 B / -9.1% (better)
reflectcall/j32-goos-js/Go 2191221 B 0 B / +0.0% 0 B 0 B / 0.0%
reflectcall/j32-goos-js/LLGo 1598592 B +67613 B / +4.4% (worse) 79769 B -9974 B / -11.1% (better)
reflectcall/j64-emscripten-memory64/LLGo 1426756 B +53508 B / +3.9% (worse) 99026 B -10466 B / -9.6% (better)
reflectcall/w32-goos-wasip1/Go 2205707 B 0 B / +0.0% 0 B 0 B / 0.0%
reflectcall/w32-goos-wasip1/LLGo 1284659 B +764 B / +0.1% (worse) 0 B 0 B / 0.0%
reflectcall/w32-wasi/LLGo 1282204 B +784 B / +0.1% (worse) 0 B 0 B / 0.0%
w32-goos-wasip1/Go 1909947 B 0 B / +0.0% 0 B 0 B / 0.0%
w32-goos-wasip1/LLGo 154611 B +668 B / +0.4% (worse) 0 B 0 B / 0.0%
w32-wasi/LLGo 155314 B +688 B / +0.4% (worse) 0 B 0 B / 0.0%
LLGo WebAssembly build measurements
Example and profile Build vs base
j32-emscripten 6.958 s -175.3 ms / -2.5% (better)
j32-goos-js 6.773 s -219.4 ms / -3.1% (better)
j64-emscripten-memory64 6.227 s -45.13 ms / -0.7% (better)
reflectcall/w32-wasi 23.657 s -501 ms / -2.1% (better)
w32-goos-wasip1 4.643 s +143 ms / +3.2% (worse)
w32-wasi 4.227 s -761.7 us / -0.01802% (better)

Compared with be669acd80cc measured in the same runner job.

@github-actions

github-actions Bot commented Oct 6, 2026 •

Copy link
Copy Markdown

LLGo baseline benchmarks

218907942ded | workflow run | long-term charts

Program measurements

Platform Workload File size vs base Text size vs base Build vs base Run vs base
Linux cprintf 8568 B 0 B / +0.0% 387 B 0 B / +0.0% 1.262 s -94.29 ms / -7.0% (better) 1.505 ms +139.1 us / +10.2% (worse)
Linux cprintf-lto 8408 B 0 B / +0.0% 368 B 0 B / +0.0% 1.278 s -91.47 ms / -6.7% (better) 1.677 ms +277.3 us / +19.8% (worse)
Linux fmtprintf 4886288 B 0 B / +0.0% 504761 B 0 B / +0.0% 9.299 s -204.4 ms / -2.2% (better) 3.505 ms +2.177 us / +0.1% (worse)
Linux fmtprintf-lto 3616568 B 0 B / +0.0% 443075 B 0 B / +0.0% 19.567 s -630.5 ms / -3.1% (better) 3.347 ms -164 us / -4.7% (better)
Linux println 675040 B 0 B / +0.0% 16855 B 0 B / +0.0% 1.299 s -34.29 ms / -2.6% (better) 1.822 ms -2.433 us / -0.1% (better)
Linux println-lto 190040 B 0 B / +0.0% 14273 B 0 B / +0.0% 1.538 s -169.3 ms / -9.9% (better) 1.710 ms -193.3 us / -10.2% (better)
macOS cprintf 50736 B 0 B / +0.0% 4409 B 0 B / +0.0% 1.312 s -427.8 ms / -24.6% (better) 5.033 ms -669.9 us / -11.7% (better)
macOS cprintf-lto 50496 B 0 B / +0.0% 161 B 0 B / +0.0% 1.352 s -372.5 ms / -21.6% (better) 4.475 ms -1.774 ms / -28.4% (better)
macOS fmtprintf 1774368 B 0 B / +0.0% 880529 B 0 B / +0.0% 8.701 s +138.1 ms / +1.6% (worse) 13.852 ms +237.5 us / +1.7% (worse)
macOS fmtprintf-lto 1361232 B 0 B / +0.0% 764957 B 0 B / +0.0% 15.387 s -3.187 s / -17.2% (better) 6.116 ms -352.8 us / -5.5% (better)
macOS println 99344 B 0 B / +0.0% 24216 B 0 B / +0.0% 1.207 s -304.9 ms / -20.2% (better) 5.420 ms +107 us / +2.0% (worse)
macOS println-lto 83664 B 0 B / +0.0% 21457 B 0 B / +0.0% 1.883 s +178.3 ms / +10.5% (worse) 7.335 ms +918.3 us / +14.3% (worse)
Windows MinGW cprintf 651776 B 0 B / +0.0% 4550 B 0 B / +0.0% 1.159 s -127.6 ms / -9.9% (better) 2.254 ms -553.8 us / -19.7% (better)
Windows MinGW cprintf-lto 43520 B 0 B / +0.0% 4486 B 0 B / +0.0% 1.069 s -241.5 ms / -18.4% (better) 2.219 ms -58.2 us / -2.6% (better)
Windows MinGW fmtprintf 5440000 B 0 B / +0.0% 604454 B 0 B / +0.0% 4.701 s -524.5 ms / -10.0% (better) 5.345 ms -618.6 us / -10.4% (better)
Windows MinGW fmtprintf-lto 4144128 B 0 B / +0.0% 551574 B 0 B / +0.0% 9.807 s -1.221 s / -11.1% (better) 5.363 ms -267.7 us / -4.8% (better)
Windows MinGW println 707072 B 0 B / +0.0% 25190 B 0 B / +0.0% 1.072 s -214.9 ms / -16.7% (better) 4.336 ms -771.4 us / -15.1% (better)
Windows MinGW println-lto 208896 B 0 B / +0.0% 22054 B 0 B / +0.0% 1.260 s -242.8 ms / -16.2% (better) 4.323 ms -783.4 us / -15.3% (better)
Windows MinGW 386 cprintf 601600 B 0 B / +0.0% 5326 B 0 B / +0.0% 1.803 s -4.385 ms / -0.2% (better) 6.051 ms +32.7 us / +0.5% (worse)
Windows MinGW 386 cprintf-lto 103424 B 0 B / +0.0% 5094 B 0 B / +0.0% 1.843 s +44.82 ms / +2.5% (worse) 5.744 ms +561.4 us / +10.8% (worse)
Windows MinGW 386 fmtprintf 4745728 B 0 B / +0.0% 472478 B 0 B / +0.0% 8.229 s +562.9 ms / +7.3% (worse) 13.145 ms +2.47 ms / +23.1% (worse)
Windows MinGW 386 fmtprintf-lto 4148736 B 0 B / +0.0% 451114 B 0 B / +0.0% 15.571 s +529.6 ms / +3.5% (worse) 10.962 ms -943.7 us / -7.9% (better)
Windows MinGW 386 println 653312 B 0 B / +0.0% 21490 B 0 B / +0.0% 1.829 s +3.4 ms / +0.2% (worse) 9.894 ms +1.08 ms / +12.3% (worse)
Windows MinGW 386 println-lto 258560 B 0 B / +0.0% 19306 B 0 B / +0.0% 2.153 s +131.2 ms / +6.5% (worse) 8.983 ms +82.5 us / +0.9% (worse)
Windows MinGW ARM64 cprintf 660992 B 0 B / +0.0% 4408 B 0 B / +0.0% 2.109 s +29.49 ms / +1.4% (worse) 7.763 ms +270.7 us / +3.6% (worse)
Windows MinGW ARM64 cprintf-lto 43520 B 0 B / +0.0% 4340 B 0 B / +0.0% 2.102 s +1.127 ms / +0.1% (worse) 7.437 ms -163.9 us / -2.2% (better)
Windows MinGW ARM64 fmtprintf 5348864 B 0 B / +0.0% 512636 B 0 B / +0.0% 7.660 s -42.9 ms / -0.6% (better) 16.195 ms +1.41 ms / +9.5% (worse)
Windows MinGW ARM64 fmtprintf-lto 4312064 B 0 B / +0.0% 478828 B 0 B / +0.0% 14.951 s -11.42 ms / -0.1% (better) 15.496 ms +623.5 us / +4.2% (worse)
Windows MinGW ARM64 println 714240 B 0 B / +0.0% 23884 B 0 B / +0.0% 2.113 s +47.46 ms / +2.3% (worse) 13.461 ms +585.5 us / +4.5% (worse)
Windows MinGW ARM64 println-lto 215552 B 0 B / +0.0% 21232 B 0 B / +0.0% 2.399 s +49.09 ms / +2.1% (worse) 12.952 ms +903 us / +7.5% (worse)
Windows MSVC cprintf 893952 B 0 B / +0.0% 65798 B 0 B / +0.0% 1.585 s +51.46 ms / +3.4% (worse) 3.740 ms -47.5 us / -1.3% (better)
Windows MSVC cprintf-lto 289792 B 0 B / +0.0% 65734 B 0 B / +0.0% 1.602 s +16.74 ms / +1.1% (worse) 5.166 ms +1.609 ms / +45.2% (worse)
Windows MSVC fmtprintf 5741056 B 0 B / +0.0% 700022 B 0 B / +0.0% 7.108 s +66.27 ms / +0.9% (worse) 10.066 ms -877.1 us / -8.0% (better)
Windows MSVC fmtprintf-lto 4467200 B 0 B / +0.0% 651110 B 0 B / +0.0% 14.194 s -68.74 ms / -0.5% (better) 10.105 ms -474.8 us / -4.5% (better)
Windows MSVC println 1017856 B 0 B / +0.0% 120854 B 0 B / +0.0% 1.562 s +5.575 ms / +0.4% (worse) 7.888 ms -160.5 us / -2.0% (better)
Windows MSVC println-lto 528384 B 0 B / +0.0% 118390 B 0 B / +0.0% 1.834 s +956.2 us / +0.1% (worse) 7.875 ms -55.3 us / -0.7% (better)
Windows MSVC 386 cprintf 513536 B 0 B / +0.0% 3931 B 0 B / +0.0% 1.592 s +27.43 ms / +1.8% (worse) 5.498 ms -1.053 ms / -16.1% (better)
Windows MSVC 386 cprintf-lto 44032 B 0 B / +0.0% 3853 B 0 B / +0.0% 1.527 s -40.8 ms / -2.6% (better) 6.058 ms -404.9 us / -6.3% (better)
Windows MSVC 386 fmtprintf 4480000 B 0 B / +0.0% 455868 B 0 B / +0.0% 7.214 s -41.54 ms / -0.6% (better) 14.025 ms +2.787 ms / +24.8% (worse)
Windows MSVC 386 fmtprintf-lto 3899904 B 0 B / +0.0% 426651 B 0 B / +0.0% 13.537 s +43.51 ms / +0.3% (worse) 11.146 ms -1.516 ms / -12.0% (better)
Windows MSVC 386 println 567808 B 0 B / +0.0% 20340 B 0 B / +0.0% 1.509 s -269.9 ms / -15.2% (better) 9.337 ms -1.348 ms / -12.6% (better)
Windows MSVC 386 println-lto 199168 B 0 B / +0.0% 18501 B 0 B / +0.0% 1.766 s -75.07 ms / -4.1% (better) 9.442 ms -665.5 us / -6.6% (better)
Windows MSVC ARM64 cprintf 662528 B 0 B / +0.0% 4192 B 0 B / +0.0% 1.612 s +25.94 ms / +1.6% (worse) 6.403 ms -487.5 us / -7.1% (better)
Windows MSVC ARM64 cprintf-lto 47616 B 0 B / +0.0% 4084 B 0 B / +0.0% 1.602 s +24.78 ms / +1.6% (worse) 6.948 ms +219.5 us / +3.3% (worse)
Windows MSVC ARM64 fmtprintf 5345280 B 0 B / +0.0% 512580 B 0 B / +0.0% 6.736 s +174.9 ms / +2.7% (worse) 14.398 ms +792 us / +5.8% (worse)
Windows MSVC ARM64 fmtprintf-lto 4318208 B 0 B / +0.0% 479488 B 0 B / +0.0% 12.872 s +88.59 ms / +0.7% (worse) 14.631 ms -212.7 us / -1.4% (better)
Windows MSVC ARM64 println 715776 B 0 B / +0.0% 23908 B 0 B / +0.0% 1.598 s +13.83 ms / +0.9% (worse) 11.726 ms +233.4 us / +2.0% (worse)
Windows MSVC ARM64 println-lto 220672 B 0 B / +0.0% 21380 B 0 B / +0.0% 1.826 s +4.531 ms / +0.2% (worse) 11.681 ms -574 us / -4.7% (better)
Core language and compiler benchmarks
Platform Benchmark ns/op vs base
Linux BenchmarkLookupPCRandom 14.760 ns/op 0 ns/op / +0.0%
Linux BenchmarkMergeCompilerFlags 207.600 ns/op +4.7 ns/op / +2.3% (worse)
Linux BenchmarkMergeLinkerFlags 137.700 ns/op -2.9 ns/op / -2.1% (better)
Linux BenchmarkChannelBuffered 55.500 ns/op +0.6 ns/op / +1.1% (worse)
Linux BenchmarkChannelHandoff 13336 ns/op +774 ns/op / +6.2% (worse)
Linux BenchmarkDefer 50.060 ns/op -2.35 ns/op / -4.5% (better)
Linux BenchmarkDirectCall 1.566 ns/op +0.053 ns/op / +3.5% (worse)
Linux BenchmarkGlobalRead 1.167 ns/op -0.392 ns/op / -25.1% (better)
Linux BenchmarkGlobalWrite 7.762 ns/op 0 ns/op / +0.0%
Linux BenchmarkGoroutine 37099 ns/op +10876 ns/op / +41.5% (worse)
Linux BenchmarkInterfaceCall 6.191 ns/op -0.022 ns/op / -0.4% (better)
Linux BenchmarkRuntimeGetG 2.959 ns/op -0.08 ns/op / -2.6% (better)
macOS BenchmarkLookupPCRandom 18.050 ns/op -0.82 ns/op / -4.3% (better)
macOS BenchmarkMergeCompilerFlags 168.100 ns/op -23.6 ns/op / -12.3% (better)
macOS BenchmarkMergeLinkerFlags 103.300 ns/op -31 ns/op / -23.1% (better)
macOS BenchmarkChannelBuffered 27.560 ns/op -9.39 ns/op / -25.4% (better)
macOS BenchmarkChannelHandoff 9813 ns/op -2070 ns/op / -17.4% (better)
macOS BenchmarkDefer 39.500 ns/op -10.34 ns/op / -20.7% (better)
macOS BenchmarkDirectCall 1.284 ns/op -0.003 ns/op / -0.2% (better)
macOS BenchmarkGlobalRead 1.416 ns/op +0.12 ns/op / +9.3% (worse)
macOS BenchmarkGlobalWrite 1.780 ns/op +0.222 ns/op / +14.2% (worse)
macOS BenchmarkGoroutine 68663 ns/op +5585 ns/op / +8.9% (worse)
macOS BenchmarkInterfaceCall 4.869 ns/op -0.661 ns/op / -12.0% (better)
macOS BenchmarkRuntimeGetG 2.885 ns/op +0.263 ns/op / +10.0% (worse)
Windows MinGW BenchmarkLookupPCRandom 7.255 ns/op -0.053 ns/op / -0.7% (better)
Windows MinGW BenchmarkMergeCompilerFlags 282.200 ns/op -4.1 ns/op / -1.4% (better)
Windows MinGW BenchmarkMergeLinkerFlags 259.600 ns/op -1.2 ns/op / -0.5% (better)
Windows MinGW BenchmarkChannelBuffered 28.400 ns/op -1.03 ns/op / -3.5% (better)
Windows MinGW BenchmarkChannelHandoff 837 ns/op +77.1 ns/op / +10.1% (worse)
Windows MinGW BenchmarkDefer 33.870 ns/op -3.5 ns/op / -9.4% (better)
Windows MinGW BenchmarkDirectCall 0.888 ns/op -0.0354 ns/op / -3.8% (better)
Windows MinGW BenchmarkGlobalRead 0.888 ns/op -0.0369 ns/op / -4.0% (better)
Windows MinGW BenchmarkGlobalWrite 4.497 ns/op -0.124 ns/op / -2.7% (better)
Windows MinGW BenchmarkGoroutine 43284 ns/op +1497 ns/op / +3.6% (worse)
Windows MinGW BenchmarkInterfaceCall 4.441 ns/op -0.256 ns/op / -5.5% (better)
Windows MinGW BenchmarkRuntimeGetG 1.076 ns/op +0.034 ns/op / +3.3% (worse)
Windows MinGW 386 BenchmarkLookupPCRandom 26.590 ns/op +0.07 ns/op / +0.3% (worse)
Windows MinGW 386 BenchmarkMergeCompilerFlags 773 ns/op +3.7 ns/op / +0.5% (worse)
Windows MinGW 386 BenchmarkMergeLinkerFlags 729.400 ns/op +32.1 ns/op / +4.6% (worse)
Windows MinGW 386 BenchmarkChannelBuffered 39.820 ns/op -1.32 ns/op / -3.2% (better)
Windows MinGW 386 BenchmarkChannelHandoff 870.900 ns/op +13.3 ns/op / +1.6% (worse)
Windows MinGW 386 BenchmarkDefer 44.080 ns/op -0.98 ns/op / -2.2% (better)
Windows MinGW 386 BenchmarkDirectCall 1.550 ns/op +0.002 ns/op / +0.1% (worse)
Windows MinGW 386 BenchmarkGlobalRead 1.550 ns/op -0.002 ns/op / -0.1% (better)
Windows MinGW 386 BenchmarkGlobalWrite 7.780 ns/op -0.004 ns/op / -0.1% (better)
Windows MinGW 386 BenchmarkGoroutine 111293 ns/op +225 ns/op / +0.2% (worse)
Windows MinGW 386 BenchmarkInterfaceCall 8.372 ns/op +0.002 ns/op / +0.02389% (worse)
Windows MinGW 386 BenchmarkRuntimeGetG 1.930 ns/op +0.001 ns/op / +0.1% (worse)
Windows MinGW ARM64 BenchmarkLookupPCRandom 12.040 ns/op 0 ns/op / +0.0%
Windows MinGW ARM64 BenchmarkMergeCompilerFlags 573.200 ns/op -10 ns/op / -1.7% (better)
Windows MinGW ARM64 BenchmarkMergeLinkerFlags 544.600 ns/op +3.9 ns/op / +0.7% (worse)
Windows MinGW ARM64 BenchmarkChannelBuffered 38.780 ns/op -0.83 ns/op / -2.1% (better)
Windows MinGW ARM64 BenchmarkChannelHandoff 3154 ns/op +320 ns/op / +11.3% (worse)
Windows MinGW ARM64 BenchmarkDefer 60.090 ns/op +0.89 ns/op / +1.5% (worse)
Windows MinGW ARM64 BenchmarkDirectCall 0.590 ns/op -0.0739 ns/op / -11.1% (better)
Windows MinGW ARM64 BenchmarkGlobalRead 0.663 ns/op -0.0007 ns/op / -0.1% (better)
Windows MinGW ARM64 BenchmarkGlobalWrite 0.737 ns/op -0.0001 ns/op / -0.01356% (better)
Windows MinGW ARM64 BenchmarkGoroutine 65851 ns/op -1143 ns/op / -1.7% (better)
Windows MinGW ARM64 BenchmarkInterfaceCall 4.151 ns/op -0.085 ns/op / -2.0% (better)
Windows MinGW ARM64 BenchmarkRuntimeGetG 1.812 ns/op +0.003 ns/op / +0.2% (worse)
Windows MSVC BenchmarkLookupPCRandom 11.680 ns/op -0.01 ns/op / -0.1% (better)
Windows MSVC BenchmarkMergeCompilerFlags 578.600 ns/op +9.6 ns/op / +1.7% (worse)
Windows MSVC BenchmarkMergeLinkerFlags 514.100 ns/op -6.9 ns/op / -1.3% (better)
Windows MSVC BenchmarkChannelBuffered 44.470 ns/op -0.73 ns/op / -1.6% (better)
Windows MSVC BenchmarkChannelHandoff 2613 ns/op +1118 ns/op / +74.8% (worse)
Windows MSVC BenchmarkDefer 69.900 ns/op +2.04 ns/op / +3.0% (worse)
Windows MSVC BenchmarkDirectCall 1.062 ns/op -0.101 ns/op / -8.7% (better)
Windows MSVC BenchmarkGlobalRead 1.146 ns/op +0.044 ns/op / +4.0% (worse)
Windows MSVC BenchmarkGlobalWrite 8.698 ns/op +0.007 ns/op / +0.1% (worse)
Windows MSVC BenchmarkGoroutine 77519 ns/op -2768 ns/op / -3.4% (better)
Windows MSVC BenchmarkInterfaceCall 6.193 ns/op -0.155 ns/op / -2.4% (better)
Windows MSVC BenchmarkRuntimeGetG 1.950 ns/op -0.031 ns/op / -1.6% (better)
Windows MSVC 386 BenchmarkLookupPCRandom 26.470 ns/op -0.23 ns/op / -0.9% (better)
Windows MSVC 386 BenchmarkMergeCompilerFlags 718.600 ns/op -42.7 ns/op / -5.6% (better)
Windows MSVC 386 BenchmarkMergeLinkerFlags 669.900 ns/op -34.6 ns/op / -4.9% (better)
Windows MSVC 386 BenchmarkChannelBuffered 38.910 ns/op +0.07 ns/op / +0.2% (worse)
Windows MSVC 386 BenchmarkChannelHandoff 905.300 ns/op +106.2 ns/op / +13.3% (worse)
Windows MSVC 386 BenchmarkDefer 49.500 ns/op +2.71 ns/op / +5.8% (worse)
Windows MSVC 386 BenchmarkDirectCall 1.866 ns/op +0.32 ns/op / +20.7% (worse)
Windows MSVC 386 BenchmarkGlobalRead 1.550 ns/op 0 ns/op / +0.0%
Windows MSVC 386 BenchmarkGlobalWrite 7.782 ns/op +0.005 ns/op / +0.1% (worse)
Windows MSVC 386 BenchmarkGoroutine 111422 ns/op +1628 ns/op / +1.5% (worse)
Windows MSVC 386 BenchmarkInterfaceCall 8.419 ns/op +0.046 ns/op / +0.5% (worse)
Windows MSVC 386 BenchmarkRuntimeGetG 1.934 ns/op -0.235 ns/op / -10.8% (better)
Windows MSVC ARM64 BenchmarkLookupPCRandom 12.210 ns/op +0.1 ns/op / +0.8% (worse)
Windows MSVC ARM64 BenchmarkMergeCompilerFlags 566.500 ns/op +3.9 ns/op / +0.7% (worse)
Windows MSVC ARM64 BenchmarkMergeLinkerFlags 526.900 ns/op -0.6 ns/op / -0.1% (better)
Windows MSVC ARM64 BenchmarkChannelBuffered 37.560 ns/op -0.04 ns/op / -0.1% (better)
Windows MSVC ARM64 BenchmarkChannelHandoff 2588 ns/op -165 ns/op / -6.0% (better)
Windows MSVC ARM64 BenchmarkDefer 62.250 ns/op +0.32 ns/op / +0.5% (worse)
Windows MSVC ARM64 BenchmarkDirectCall 0.590 ns/op -0.0735 ns/op / -11.1% (better)
Windows MSVC ARM64 BenchmarkGlobalRead 0.664 ns/op +0.0006 ns/op / +0.1% (worse)
Windows MSVC ARM64 BenchmarkGlobalWrite 3.832 ns/op +0.038 ns/op / +1.0% (worse)
Windows MSVC ARM64 BenchmarkGoroutine 57759 ns/op -2682 ns/op / -4.4% (better)
Windows MSVC ARM64 BenchmarkInterfaceCall 4.139 ns/op -0.096 ns/op / -2.3% (better)
Windows MSVC ARM64 BenchmarkRuntimeGetG 1.768 ns/op -0.001 ns/op / -0.1% (better)
Timer runtime benchmarks
Platform Operation and runtime ns/op vs base
Linux AfterFuncZeroDelivery/Go 920.500 ns/op -126.5 ns/op / -12.1% (better)
Linux AfterFuncZeroDelivery/LLGo 39230 ns/op -2788 ns/op / -6.6% (better)
Linux CreateStop/Go 306.500 ns/op +0.3 ns/op / +0.1% (worse)
Linux CreateStop/LLGo 2019 ns/op +65 ns/op / +3.3% (worse)
Linux RearmStopped/Go 115.900 ns/op +1 ns/op / +0.9% (worse)
Linux RearmStopped/LLGo 1395 ns/op +136 ns/op / +10.8% (worse)
Linux ResetActive/Go 68.650 ns/op +1.03 ns/op / +1.5% (worse)
Linux ResetActive/LLGo 822 ns/op -40 ns/op / -4.6% (better)
Linux ResetHeap1024/Go 67.100 ns/op -0.1 ns/op / -0.1% (better)
Linux ResetHeap1024/LLGo 180 ns/op -3.9 ns/op / -2.1% (better)
macOS AfterFuncZeroDelivery/Go 699.400 ns/op +146.4 ns/op / +26.5% (worse)
macOS AfterFuncZeroDelivery/LLGo 108513 ns/op +2021 ns/op / +1.9% (worse)
macOS CreateStop/Go 203 ns/op -25.6 ns/op / -11.2% (better)
macOS CreateStop/LLGo 668.100 ns/op -89.8 ns/op / -11.8% (better)
macOS RearmStopped/Go 78 ns/op -3.79 ns/op / -4.6% (better)
macOS RearmStopped/LLGo 370.300 ns/op -76.7 ns/op / -17.2% (better)
macOS ResetActive/Go 56.480 ns/op +1.72 ns/op / +3.1% (worse)
macOS ResetActive/LLGo 167.300 ns/op -58.5 ns/op / -25.9% (better)
macOS ResetHeap1024/Go 61.290 ns/op +8.31 ns/op / +15.7% (worse)
macOS ResetHeap1024/LLGo 91.810 ns/op -6.43 ns/op / -6.5% (better)
Windows MinGW AfterFuncZeroDelivery/Go 431.700 ns/op +1 ns/op / +0.2% (worse)
Windows MinGW AfterFuncZeroDelivery/LLGo 86768 ns/op -1572 ns/op / -1.8% (better)
Windows MinGW CreateStop/Go 122.500 ns/op +0.7 ns/op / +0.6% (worse)
Windows MinGW CreateStop/LLGo 338.500 ns/op +16 ns/op / +5.0% (worse)
Windows MinGW RearmStopped/Go 47.650 ns/op -0.06 ns/op / -0.1% (better)
Windows MinGW RearmStopped/LLGo 221.600 ns/op +36.5 ns/op / +19.7% (worse)
Windows MinGW ResetActive/Go 21.430 ns/op +0.09 ns/op / +0.4% (worse)
Windows MinGW ResetActive/LLGo 504.200 ns/op -42.8 ns/op / -7.8% (better)
Windows MinGW ResetHeap1024/Go 21.040 ns/op +0.24 ns/op / +1.2% (worse)
Windows MinGW ResetHeap1024/LLGo 93.480 ns/op +4.54 ns/op / +5.1% (worse)
Windows MinGW 386 AfterFuncZeroDelivery/Go 951.600 ns/op -18.8 ns/op / -1.9% (better)
Windows MinGW 386 AfterFuncZeroDelivery/LLGo 202341 ns/op -410 ns/op / -0.2% (better)
Windows MinGW 386 CreateStop/Go 199.400 ns/op +9.7 ns/op / +5.1% (worse)
Windows MinGW 386 CreateStop/LLGo 539.500 ns/op +18.4 ns/op / +3.5% (worse)
Windows MinGW 386 RearmStopped/Go 63.410 ns/op +0.26 ns/op / +0.4% (worse)
Windows MinGW 386 RearmStopped/LLGo 365.300 ns/op -31.8 ns/op / -8.0% (better)
Windows MinGW 386 ResetActive/Go 39.010 ns/op +0.06 ns/op / +0.2% (worse)
Windows MinGW 386 ResetActive/LLGo 1002 ns/op +558 ns/op / +125.7% (worse)
Windows MinGW 386 ResetHeap1024/Go 39.480 ns/op +0.16 ns/op / +0.4% (worse)
Windows MinGW 386 ResetHeap1024/LLGo 188 ns/op -0.3 ns/op / -0.2% (better)
Windows MinGW ARM64 AfterFuncZeroDelivery/Go 666.800 ns/op +2 ns/op / +0.3% (worse)
Windows MinGW ARM64 AfterFuncZeroDelivery/LLGo 155677 ns/op -1092 ns/op / -0.7% (better)
Windows MinGW ARM64 CreateStop/Go 196.100 ns/op +1.3 ns/op / +0.7% (worse)
Windows MinGW ARM64 CreateStop/LLGo 361.900 ns/op -1.2 ns/op / -0.3% (better)
Windows MinGW ARM64 RearmStopped/Go 70.590 ns/op -0.11 ns/op / -0.2% (better)
Windows MinGW ARM64 RearmStopped/LLGo 249.600 ns/op -1.1 ns/op / -0.4% (better)
Windows MinGW ARM64 ResetActive/Go 30.930 ns/op -0.13 ns/op / -0.4% (better)
Windows MinGW ARM64 ResetActive/LLGo 122.500 ns/op +0.1 ns/op / +0.1% (worse)
Windows MinGW ARM64 ResetHeap1024/Go 31.140 ns/op 0 ns/op / +0.0%
Windows MinGW ARM64 ResetHeap1024/LLGo 125.700 ns/op -0.4 ns/op / -0.3% (better)
Windows MSVC AfterFuncZeroDelivery/Go 703.900 ns/op -5.8 ns/op / -0.8% (better)
Windows MSVC AfterFuncZeroDelivery/LLGo 149185 ns/op +2063 ns/op / +1.4% (worse)
Windows MSVC CreateStop/Go 193.200 ns/op -1.1 ns/op / -0.6% (better)
Windows MSVC CreateStop/LLGo 1026 ns/op +120.8 ns/op / +13.3% (worse)
Windows MSVC RearmStopped/Go 71.750 ns/op +0.04 ns/op / +0.1% (worse)
Windows MSVC RearmStopped/LLGo 352.500 ns/op -12.4 ns/op / -3.4% (better)
Windows MSVC ResetActive/Go 31.700 ns/op -0.05 ns/op / -0.2% (better)
Windows MSVC ResetActive/LLGo 186.700 ns/op -9 ns/op / -4.6% (better)
Windows MSVC ResetHeap1024/Go 32.090 ns/op +0.11 ns/op / +0.3% (worse)
Windows MSVC ResetHeap1024/LLGo 133.700 ns/op +0.4 ns/op / +0.3% (worse)
Windows MSVC 386 AfterFuncZeroDelivery/Go 969.400 ns/op +20.4 ns/op / +2.1% (worse)
Windows MSVC 386 AfterFuncZeroDelivery/LLGo 208527 ns/op +1337 ns/op / +0.6% (worse)
Windows MSVC 386 CreateStop/Go 193.900 ns/op +2.8 ns/op / +1.5% (worse)
Windows MSVC 386 CreateStop/LLGo 493.800 ns/op +14.2 ns/op / +3.0% (worse)
Windows MSVC 386 RearmStopped/Go 63.440 ns/op +0.13 ns/op / +0.2% (worse)
Windows MSVC 386 RearmStopped/LLGo 328.200 ns/op +5.1 ns/op / +1.6% (worse)
Windows MSVC 386 ResetActive/Go 39.010 ns/op +0.04 ns/op / +0.1% (worse)
Windows MSVC 386 ResetActive/LLGo 928.800 ns/op -6.4 ns/op / -0.7% (better)
Windows MSVC 386 ResetHeap1024/Go 39.610 ns/op +0.21 ns/op / +0.5% (worse)
Windows MSVC 386 ResetHeap1024/LLGo 177.700 ns/op +6.9 ns/op / +4.0% (worse)
Windows MSVC ARM64 AfterFuncZeroDelivery/Go 666.900 ns/op +8.9 ns/op / +1.4% (worse)
Windows MSVC ARM64 AfterFuncZeroDelivery/LLGo 170514 ns/op +7929 ns/op / +4.9% (worse)
Windows MSVC ARM64 CreateStop/Go 197 ns/op -2.5 ns/op / -1.3% (better)
Windows MSVC ARM64 CreateStop/LLGo 408.600 ns/op +11.8 ns/op / +3.0% (worse)
Windows MSVC ARM64 RearmStopped/Go 70.580 ns/op +0.07 ns/op / +0.1% (worse)
Windows MSVC ARM64 RearmStopped/LLGo 272.700 ns/op +2.6 ns/op / +1.0% (worse)
Windows MSVC ARM64 ResetActive/Go 31.140 ns/op +0.05 ns/op / +0.2% (worse)
Windows MSVC ARM64 ResetActive/LLGo 132.200 ns/op -1.4 ns/op / -1.0% (better)
Windows MSVC ARM64 ResetHeap1024/Go 31.160 ns/op +0.06 ns/op / +0.2% (worse)
Windows MSVC ARM64 ResetHeap1024/LLGo 137.100 ns/op +1.6 ns/op / +1.2% (worse)

Compared with be669acd80cc measured in the same runner job.

@cpunion
cpunion force-pushed the codex/wasm-reflect-worker-tables-20261006 branch from 2861a8d to 2189079 Compare October 6, 2026 09:38
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant