Skip to content

debug: inspect parked Wasm goroutine source stacks - #2742

Open
cpunion wants to merge 3 commits into
xgo-dev:mainfrom
cpunion:codex/wasm-debug-runtime-20261006
Open

cpunion wants to merge 3 commits into
xgo-dev:mainfrom
cpunion:codex/wasm-debug-runtime-20261006

Conversation

@cpunion

@cpunion cpunion commented Oct 6, 2026 •

Copy link
Copy Markdown
Collaborator

Browser debug builds retain logical Go caller frames and a versioned runtime registry, so DevTools $goroutines can show live IDs, parent IDs, state, last processor and source stacks even when a fiber is parked. llgo debug enables the opt-in tag; ordinary builds retain their existing instrumentation policy. Decoding lives in a separately tested extension module, with batched memory reads, sequence guards and lazy snapshot objects that are released with their children.

Validation: full compiler/browser/debugger/schema Go tests; 24 extension tests; real Chrome for Testing 152 with Memory32/64, 1/2/4 workers, embedded/external DWARF, parked recursive stacks after GC, valid P0 display, and record removal after exit. Rebased onto main be669acd8; before this rebase the Wasm/browser/runtime and patch coverage checks passed with host checks pending. Current rebase CI is running; all 24 extension tests pass on this head. A local integration with #2738/#2743 passes the same ten Chrome sessions. The same browser acceptance runs in PR CI.

Users still pause all worker threads in DevTools before inspecting shared state. Logical frames expose function/file/line, not parked-frame locals or automatic worker control. Advances the browser runtime views in #2164; WASI guest-debug transport and physical-board qualification remain separate.

@cpunion

cpunion commented Oct 6, 2026

Copy link
Copy Markdown
Collaborator Author

@fennoai please review the opt-in Wasm logical goroutine registry, caller-frame instrumentation and standalone decoder, especially GC/lifetime safety, paused-worker consistency and default-build overhead. Full compiler/browser/debugger/schema tests and the actual Chrome Memory32/64 1/2/4-worker matrix pass locally, including parked recursive stacks after GC and record removal after exit. The frontend still requires pausing other worker threads before shared-memory inspection; it does not claim automatic worker control or parked-frame locals.

@fennoai fennoai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review: WASM debug runtime + goroutines panel

Thorough, well-engineered PR. The lock-free registry snapshot protocol (epoch at the registry level, per-node odd/even sequence guards with begin/end double-reads, re-verified in the JS reader after the walk) is clean and well-tested (Memory32/64, 1/2/4 workers, corruption cases). Build-tag partitioning across the stub / wasm-debug / other variants fully covers with no overlap. Docs (dev/browser-debugging.md, doc/debugging.md, the READMEs) and schema_v1.json accurately match the runtime behavior and consumer fields.

Security: no actionable issues. Goroutine strings reach DevTools as structured protocol objects ({type:'string', ...}), not HTML, so no DOM-injection sink is introduced; keep them text-only downstream. Memory reads are bounded (LIMIT = 4096, cycle detection, numberAddress range checks, Chrome's bounds-checked getWasmLinearMemory).

Findings below are minor/perf — nothing blocking.

Comment thread cmd/internal/browser/extension/goroutines.js
Comment thread cmd/internal/debug/debug.go
Comment thread runtime/internal/runtime/traceback_threads_wasm_debug.go Outdated
Comment thread runtime/internal/runtime/traceback_threads_wasm_debug.go
@fennoai

fennoai Bot commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

Focused review: GC/lifetime, paused-worker consistency, default-build overhead

I verified each of your three highlighted concerns directly against the code (locked range 78973102...9cae6c0). Short version: all three hold up. Details below.

✅ Default-build overhead — zero

wasmDebuggerEnabled is a compile-time const, not a variable. The stub (traceback_wasm_debug_stub.go) sets const wasmDebuggerEnabled = false under !llgo || !wasm || baremetal || !llgo.wasm.debugger, which is the exact negation of the wasm-debug file's llgo && wasm && !baremetal && llgo.wasm.debugger — full coverage, no overlap. So in every ordinary build the eight if wasmDebuggerEnabled { … } guards added to caller.go (Push/Pop/updateCurrentFrame/store-create) are dead-code eliminated, and the three stub shims (beginWasmCallerUpdate/endWasmCallerUpdate/publishWasmCallerStore) inline to nothing. The traceback_threads_other.go tag edit (wasm → wasm && !llgo.wasm.debugger) correctly carves out only the debugger-wasm case, so native/other builds are untouched. No per-frame cost in the default path.

✅ GC / lifetime safety

  • Node storage: nodes are c.Malloc'd libc memory (not GC-scanned), which is the right call — it avoids recursing into Go allocation while registering the initial G before its local context is installed (as the comment notes).
  • node.callers → *callerLocationStore dangling: safe. The store is a GC-heap object kept alive by the //llgointernal:gls per-goroutine callerLocationStoreCurrent, i.e. for the whole G lifetime. The node's lifetime is a subset of the G's (both set up in initG→registerTraceback, both torn down in destroyG/freeRuntimeContext→unregisterTraceback), so the store can't be collected while the node still references it. unregisterTraceback frees only the node, never the store — correct.
  • node.gp / g.tracebackThread: the opaque unsafe.Pointer in g.tracebackThread mirrors the pre-existing native variant exactly, so GC's conservative treatment of that field is unchanged by this PR.
  • Reuse / staleness: each context is freshly AllocRoot'd and Memset-zeroed (no pooling that could carry a stale tracebackThread), and unregisterTraceback is idempotent (nil-guarded, clears gp.tracebackThread before c.Free), so the destroyG + freeRuntimeContext double-call path is safe. This lines up with your "record removal after exit" test.

✅ Paused-worker consistency

The seqlock protocol is correct and defensively enforced in the standalone decoder (goroutines.js):

  • Registry epoch: writers bump it odd→even around list mutation; the decoder rejects an odd initial read and re-checks equality at the end → any concurrent register/unregister is caught.
  • Per-node sequence: writers bump it odd→even around in-place frame updates (begin/endWasmCallerUpdate, with updateDepth bracketing nested updates so one transition spans the whole nested update); the decoder rejects odd and re-checks per node → any concurrent frame mutation is caught.
  • Guards: LIMIT = 4096 on both node count and frame count, cycle detection (visited), nil-G rejection, and frame.size <= 0.

Net effect: if a worker is not actually paused, the decoder fails loudly with a "pause all workers" error rather than returning torn/corrupt data — which matches your stated contract that the frontend requires pausing other workers and makes no automatic-worker-control claim.

Notes (non-blocking, carried over from the diff review)

  • goroutines.js does one awaited getWasmLinearMemory per field; per node that's ~frames * 7 serialized round-trips (up to LIMIT nodes/depth). Correct, but batching a whole node/frame-struct read would cut transport calls ~10×.
  • traceback_threads_wasm_debug.go: registerTraceback's parent param is unused on wasm (parent linkage is reconstructed in JS via parentGoid); and the return after fatal(...) is dead in the targeted js && wasm build where fatal→ForceExit is noreturn. Both cosmetic.
  • cmd/internal/debug/debug.go: llgo.wasm.debugger is appended unconditionally for the browser backend, so a user who already passed -tags=llgo.wasm.debugger gets it twice (harmless; a strings.Contains guard would be tidier).

No blocking issues on any of the three areas you flagged.

@codecov

codecov Bot commented Oct 6, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@cpunion

cpunion commented Oct 6, 2026

Copy link
Copy Markdown
Collaborator Author

Current doc_verify failure is the unchanged README LLDB homepage (https://lldb.llvm.org/): lychee reports a connection failure, with 95 other links passing and no excluded links. Direct access currently returns HTTP 200 and the official documentation remains valid. I have not retried CI, added link exclusions/tolerance, or changed a valid URL to bypass the failure. The full local Chrome debugger matrix and all 23 extension tests pass; the fresh hosted runtime/debugger checks are still running.

@cpunion

cpunion commented Oct 6, 2026

Copy link
Copy Markdown
Collaborator Author

@fennoai please check the small P0 correction in eed7d78. Runtime P IDs are zero-based; debugger records now initialize last-P to -1 and the decoder reads the signed field so P0 remains valid. A new Memory32/64 unit test covers zero and the unstarted sentinel, and the real Chrome matrix now asserts the main G is P0 and both parked Gs have valid last-P values. All 24 extension tests and the complete actual 1/2/4-worker, embedded/external matrix pass. The earlier external-link failure was left intact; this head changes runtime decoding and tests, with no URL or CI-rule changes.

@fennoai

fennoai Bot commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

P0 correction verified ✅ — zero-based processor IDs

Checked eed7d788a against the code (and the sibling 0409080d6 that also landed since my last review). The fix is correct and complete.

The P0 fix (eed7d788a) is right

  • Runtime (traceback_threads_wasm_debug.go): processorID now initializes to -1 in registerTraceback (*node = wasmDebuggerNode{gp: gp, processorID: -1}), and beginWasmCallerUpdate overwrites it with the real zero-based gp.m.p.id once the G first runs. So the sentinel is strictly -1 = never run, >= 0 = last P.
  • Decoder (goroutines.js): reads the field width-correctly as signed — BigInt.asIntN(nodeFields[spec.processor_id].size * 8, unsigned(...)) — then treats processor >= 0n as a real P (so P0 → "0") and negative as null. This is robust regardless of the DWARF signed/unsigned labelling because it keys only off the byte size. Exactly the fix for the processorID == 0 nuance from the prior review.
  • Tests: the new goroutines_test.js case asserts P0 → '0' and 0xffffffff → null across both memory widths (4 and 8); chrome_session.mjs now asserts the main G is P0, both parked Gs have non-null last-P, and (multi-worker) that they differ. Correct coverage of the zero and unstarted-sentinel edges.

Also landed since last review (0409080d6) — spot-checked, no regressions

This commit resolved all three non-blocking notes from my earlier review, and I verified it doesn't weaken the consistency guarantees:

  • Batched memory reads: the decoder now fetches each fixed struct / the frame array in one getWasmLinearMemory call and decodes scalars locally — the ~10× round-trip reduction I suggested. Crucially, the seqlock re-verification stays fresh: the trailing per-node sequence and final epoch re-checks use plugin.readNamedUnsigned (a live transport read), not the cached nodeBytes/registryBytes — so the paused-worker torn-read detection still works. The batched frame read is bounded (count <= LIMIT, memory() caps size <= LIMIT*256), well clear of realistic frame sizes.
  • Duplicate build tag: now guarded by slices.Contains(strings.Fields(strings.ReplaceAll(conf.Tags, ",", " ")), "llgo.wasm.debugger"), correctly handling comma- and space-separated tag lists.
  • Unused parent param: renamed to _ with a one-line comment.

Checks I ran

  • node --test in the extension dir: 24/24 pass.
  • gofmt -l on the changed Go files: clean.

No issues. The P0 correction does what it says, and the batching change is consistency-safe. The pre-existing external-link failure is unrelated to this head (runtime decoding + tests only, no URL/CI changes), as you noted.

@github-actions

github-actions Bot commented Oct 6, 2026 •

Copy link
Copy Markdown

LLGo WebAssembly build benchmarks

49441f338fce | workflow run | long-term charts

WebAssembly output sizes
Example, profile and compiler Wasm module vs base Generated JS glue vs base
cprintf/j32-emscripten/LLGo 149869 B +629 B / +0.4% (worse) 88157 B 0 B / +0.0%
cprintf/j32-goos-js/LLGo 149147 B +636 B / +0.4% (worse) 72985 B 0 B / +0.0%
cprintf/j64-emscripten-memory64/LLGo 137257 B +550 B / +0.4% (worse) 91328 B 0 B / +0.0%
cprintf/w32-goos-wasip1/LLGo 154895 B +599 B / +0.4% (worse) 0 B 0 B / 0.0%
cprintf/w32-wasi/LLGo 155578 B +599 B / +0.4% (worse) 0 B 0 B / 0.0%
fmtprintf/j32-emscripten/LLGo 3057896 B +897 B / +0.02934% (worse) 130373 B 0 B / +0.0%
fmtprintf/j32-goos-js/Go 2526852 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/j32-goos-js/LLGo 3153841 B +862 B / +0.02734% (worse) 100773 B 0 B / +0.0%
fmtprintf/j64-emscripten-memory64/LLGo 2819902 B +810 B / +0.02873% (worse) 135853 B 0 B / +0.0%
fmtprintf/w32-goos-wasip1/Go 2500019 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/w32-goos-wasip1/LLGo 2356287 B +839 B / +0.03562% (worse) 0 B 0 B / 0.0%
fmtprintf/w32-wasi/LLGo 2352865 B +839 B / +0.03567% (worse) 0 B 0 B / 0.0%
j32-emscripten/LLGo 149209 B +627 B / +0.4% (worse) 88157 B 0 B / +0.0%
j32-goos-js/Go 1895533 B 0 B / +0.0% 0 B 0 B / 0.0%
j32-goos-js/LLGo 148628 B +626 B / +0.4% (worse) 72985 B 0 B / +0.0%
j64-emscripten-memory64/LLGo 136598 B +551 B / +0.4% (worse) 91328 B 0 B / +0.0%
reflectcall/j32-emscripten/LLGo 1481983 B +860 B / +0.1% (worse) 104917 B 0 B / +0.0%
reflectcall/j32-goos-js/Go 2191221 B 0 B / +0.0% 0 B 0 B / 0.0%
reflectcall/j32-goos-js/LLGo 1531856 B +877 B / +0.1% (worse) 89743 B 0 B / +0.0%
reflectcall/j64-emscripten-memory64/LLGo 1374047 B +799 B / +0.1% (worse) 109492 B 0 B / +0.0%
reflectcall/w32-goos-wasip1/Go 2205707 B 0 B / +0.0% 0 B 0 B / 0.0%
reflectcall/w32-goos-wasip1/LLGo 1284734 B +839 B / +0.1% (worse) 0 B 0 B / 0.0%
reflectcall/w32-wasi/LLGo 1282259 B +839 B / +0.1% (worse) 0 B 0 B / 0.0%
w32-goos-wasip1/Go 1909947 B 0 B / +0.0% 0 B 0 B / 0.0%
w32-goos-wasip1/LLGo 154542 B +599 B / +0.4% (worse) 0 B 0 B / 0.0%
w32-wasi/LLGo 155225 B +599 B / +0.4% (worse) 0 B 0 B / 0.0%
LLGo WebAssembly build measurements
Example and profile Build vs base
j32-emscripten 6.747 s -249.9 ms / -3.6% (better)
j32-goos-js 6.831 s +150.3 ms / +2.2% (worse)
j64-emscripten-memory64 6.162 s +63.64 ms / +1.0% (worse)
reflectcall/w32-wasi 23.792 s -162.6 ms / -0.7% (better)
w32-goos-wasip1 4.330 s -110.3 ms / -2.5% (better)
w32-wasi 4.040 s -58.23 ms / -1.4% (better)

Compared with be669acd80cc measured in the same runner job.

@github-actions

github-actions Bot commented Oct 6, 2026

Copy link
Copy Markdown

LLGo baseline benchmarks

eed7d788a55b | workflow run | long-term charts

Program measurements

Platform Workload File size vs base Text size vs base Build vs base Run vs base
Linux cprintf 8568 B 0 B / +0.0% 387 B 0 B / +0.0% 1.007 s +17.6 ms / +1.8% (worse) 1.280 ms +15.39 us / +1.2% (worse)
Linux cprintf-lto 8408 B 0 B / +0.0% 368 B 0 B / +0.0% 971.589 ms +1.207 ms / +0.1% (worse) 1.246 ms +8.807 us / +0.7% (worse)
Linux fmtprintf 4882032 B +1080 B / +0.02213% (worse) 501619 B 0 B / +0.0% 7.301 s -267.7 ms / -3.5% (better) 3.221 ms +207.4 us / +6.9% (worse)
Linux fmtprintf-lto 3601056 B +248 B / +0.006887% (worse) 438538 B 0 B / +0.0% 16.721 s -414.6 ms / -2.4% (better) 2.919 ms -260.7 us / -8.2% (better)
Linux println 675832 B +880 B / +0.1% (worse) 16855 B 0 B / +0.0% 1.020 s -4.536 ms / -0.4% (better) 1.679 ms +73.84 us / +4.6% (worse)
Linux println-lto 190072 B +48 B / +0.02526% (worse) 14273 B 0 B / +0.0% 1.296 s -10.07 ms / -0.8% (better) 1.580 ms -11.72 us / -0.7% (better)
macOS cprintf 50736 B 0 B / +0.0% 4409 B 0 B / +0.0% 1.323 s -18.3 ms / -1.4% (better) 4.092 ms -6.655 ms / -61.9% (better)
macOS cprintf-lto 50496 B 0 B / +0.0% 161 B 0 B / +0.0% 1.302 s +216.9 ms / +20.0% (worse) 7.994 ms +4.579 ms / +134.1% (worse)
macOS fmtprintf 1773456 B 0 B / +0.0% 878020 B +240 B / +0.02734% (worse) 4.051 s -1.799 s / -30.8% (better) 9.993 ms -1.933 ms / -16.2% (better)
macOS fmtprintf-lto 1360928 B 0 B / +0.0% 762632 B +204 B / +0.02676% (worse) 10.216 s -995.4 ms / -8.9% (better) 4.891 ms +798.8 us / +19.5% (worse)
macOS println 99344 B 0 B / +0.0% 24216 B 0 B / +0.0% 1.478 s +555.4 ms / +60.2% (worse) 5.123 ms +867.3 us / +20.4% (worse)
macOS println-lto 83664 B 0 B / +0.0% 21457 B 0 B / +0.0% 1.077 s -152.1 ms / -12.4% (better) 3.110 ms -685.3 us / -18.1% (better)
Windows MinGW cprintf 651776 B +512 B / +0.1% (worse) 4550 B 0 B / +0.0% 1.794 s +80.91 ms / +4.7% (worse) 4.205 ms +764.8 us / +22.2% (worse)
Windows MinGW cprintf-lto 43520 B 0 B / +0.0% 4486 B 0 B / +0.0% 1.826 s +85.28 ms / +4.9% (worse) 3.829 ms +449.8 us / +13.3% (worse)
Windows MinGW fmtprintf 5434368 B +1536 B / +0.02827% (worse) 600310 B 0 B / +0.0% 7.620 s +257.2 ms / +3.5% (worse) 8.437 ms +146.5 us / +1.8% (worse)
Windows MinGW fmtprintf-lto 4127232 B 0 B / +0.0% 546582 B 0 B / +0.0% 15.233 s +250.6 ms / +1.7% (worse) 7.968 ms -438.4 us / -5.2% (better)
Windows MinGW println 707072 B +512 B / +0.1% (worse) 25190 B 0 B / +0.0% 1.781 s +96 ms / +5.7% (worse) 6.917 ms +78.5 us / +1.1% (worse)
Windows MinGW println-lto 208896 B 0 B / +0.0% 22054 B 0 B / +0.0% 2.045 s +41.27 ms / +2.1% (worse) 6.906 ms +503 us / +7.9% (worse)
Windows MinGW 386 cprintf 602624 B +1024 B / +0.2% (worse) 5326 B 0 B / +0.0% 1.276 s -6.419 ms / -0.5% (better) 3.839 ms +23.7 us / +0.6% (worse)
Windows MinGW 386 cprintf-lto 103424 B 0 B / +0.0% 5094 B 0 B / +0.0% 1.281 s -41.91 ms / -3.2% (better) 3.801 ms -65.5 us / -1.7% (better)
Windows MinGW 386 fmtprintf 4746240 B +1024 B / +0.02158% (worse) 472478 B 0 B / +0.0% 5.729 s -245.6 ms / -4.1% (better) 7.832 ms -1.309 ms / -14.3% (better)
Windows MinGW 386 fmtprintf-lto 4148736 B 0 B / +0.0% 451114 B 0 B / +0.0% 11.370 s -651.1 ms / -5.4% (better) 8.098 ms -332.5 us / -3.9% (better)
Windows MinGW 386 println 654336 B +1024 B / +0.2% (worse) 21490 B 0 B / +0.0% 1.267 s -12.11 ms / -0.9% (better) 6.606 ms -84 us / -1.3% (better)
Windows MinGW 386 println-lto 258560 B 0 B / +0.0% 19306 B 0 B / +0.0% 1.512 s -7.571 ms / -0.5% (better) 6.782 ms +61.7 us / +0.9% (worse)
Windows MinGW ARM64 cprintf 662016 B +512 B / +0.1% (worse) 4408 B 0 B / +0.0% 1.765 s +28.34 ms / +1.6% (worse) 5.568 ms -227.4 us / -3.9% (better)
Windows MinGW ARM64 cprintf-lto 43520 B 0 B / +0.0% 4340 B 0 B / +0.0% 1.805 s +39.22 ms / +2.2% (worse) 5.598 ms -123.3 us / -2.2% (better)
Windows MinGW ARM64 fmtprintf 5345792 B +2048 B / +0.03833% (worse) 510876 B 0 B / +0.0% 6.652 s +52.04 ms / +0.8% (worse) 11.274 ms +111.1 us / +1.0% (worse)
Windows MinGW ARM64 fmtprintf-lto 4302336 B 0 B / +0.0% 477276 B +12 B / +0.002514% (worse) 12.974 s +142.4 ms / +1.1% (worse) 12.349 ms +622.6 us / +5.3% (worse)
Windows MinGW ARM64 println 714752 B +512 B / +0.1% (worse) 23884 B 0 B / +0.0% 1.784 s +65.23 ms / +3.8% (worse) 9.899 ms +120.7 us / +1.2% (worse)
Windows MinGW ARM64 println-lto 215552 B 0 B / +0.0% 21232 B 0 B / +0.0% 2.011 s +40.28 ms / +2.0% (worse) 10.153 ms +241.1 us / +2.4% (worse)
Windows MSVC cprintf 893952 B 0 B / +0.0% 65798 B 0 B / +0.0% 1.128 s -160.7 ms / -12.5% (better) 2.256 ms -170.1 us / -7.0% (better)
Windows MSVC cprintf-lto 289792 B 0 B / +0.0% 65734 B 0 B / +0.0% 1.194 s -38.87 ms / -3.2% (better) 2.238 ms +113.7 us / +5.4% (worse)
Windows MSVC fmtprintf 5735424 B +2560 B / +0.04465% (worse) 695862 B 0 B / +0.0% 6.257 s +125.4 ms / +2.0% (worse) 6.238 ms -1.411 ms / -18.4% (better)
Windows MSVC fmtprintf-lto 4450816 B +1024 B / +0.02301% (worse) 646134 B 0 B / +0.0% 11.801 s +517.2 ms / +4.6% (worse) 6.454 ms +505.4 us / +8.5% (worse)
Windows MSVC println 1018368 B +1024 B / +0.1% (worse) 120854 B 0 B / +0.0% 1.313 s +41.95 ms / +3.3% (worse) 4.954 ms -1.275 ms / -20.5% (better)
Windows MSVC println-lto 528384 B 0 B / +0.0% 118390 B 0 B / +0.0% 1.327 s -201.6 ms / -13.2% (better) 4.532 ms -466.1 us / -9.3% (better)
Windows MSVC 386 cprintf 514048 B +512 B / +0.1% (worse) 3931 B 0 B / +0.0% 1.575 s +51.51 ms / +3.4% (worse) 6.061 ms +534.5 us / +9.7% (worse)
Windows MSVC 386 cprintf-lto 44032 B 0 B / +0.0% 3853 B 0 B / +0.0% 1.571 s +5.318 ms / +0.3% (worse) 5.742 ms -1.735 ms / -23.2% (better)
Windows MSVC 386 fmtprintf 4480000 B +512 B / +0.01143% (worse) 455868 B 0 B / +0.0% 7.288 s +259.4 ms / +3.7% (worse) 11.771 ms +724.9 us / +6.6% (worse)
Windows MSVC 386 fmtprintf-lto 3899904 B +512 B / +0.01313% (worse) 426651 B 0 B / +0.0% 13.426 s +111.2 ms / +0.8% (worse) 11.029 ms -379.4 us / -3.3% (better)
Windows MSVC 386 println 568320 B +1024 B / +0.2% (worse) 20340 B 0 B / +0.0% 1.536 s +20.83 ms / +1.4% (worse) 9.276 ms +35.9 us / +0.4% (worse)
Windows MSVC 386 println-lto 199168 B 0 B / +0.0% 18501 B 0 B / +0.0% 1.796 s +52.51 ms / +3.0% (worse) 9.564 ms -88.3 us / -0.9% (better)
Windows MSVC ARM64 cprintf 662528 B 0 B / +0.0% 4192 B 0 B / +0.0% 1.741 s -797 us / -0.04576% (better) 8.011 ms +451.3 us / +6.0% (worse)
Windows MSVC ARM64 cprintf-lto 47616 B 0 B / +0.0% 4084 B 0 B / +0.0% 1.762 s +30.43 ms / +1.8% (worse) 7.947 ms +149.6 us / +1.9% (worse)
Windows MSVC ARM64 fmtprintf 5341696 B +2048 B / +0.03835% (worse) 510808 B 0 B / +0.0% 7.080 s +15.69 ms / +0.2% (worse) 15.777 ms +205.5 us / +1.3% (worse)
Windows MSVC ARM64 fmtprintf-lto 4309504 B 0 B / +0.0% 477940 B +16 B / +0.003348% (worse) 13.549 s -41.33 ms / -0.3% (better) 15.558 ms -348.3 us / -2.2% (better)
Windows MSVC ARM64 println 716288 B +1024 B / +0.1% (worse) 23908 B 0 B / +0.0% 1.736 s +23.83 ms / +1.4% (worse) 14.527 ms +782.6 us / +5.7% (worse)
Windows MSVC ARM64 println-lto 220672 B 0 B / +0.0% 21380 B 0 B / +0.0% 1.984 s +28.8 ms / +1.5% (worse) 13.853 ms +241.9 us / +1.8% (worse)
Core language and compiler benchmarks
Platform Benchmark ns/op vs base
Linux BenchmarkLookupPCRandom 14.500 ns/op -0.02 ns/op / -0.1% (better)
Linux BenchmarkMergeCompilerFlags 201.800 ns/op +5.6 ns/op / +2.9% (worse)
Linux BenchmarkMergeLinkerFlags 126.800 ns/op +1.2 ns/op / +1.0% (worse)
Linux BenchmarkChannelBuffered 55.840 ns/op +0.42 ns/op / +0.8% (worse)
Linux BenchmarkChannelHandoff 14306 ns/op +778 ns/op / +5.8% (worse)
Linux BenchmarkDefer 51.180 ns/op -1.76 ns/op / -3.3% (better)
Linux BenchmarkDirectCall 1.601 ns/op +0.436 ns/op / +37.4% (worse)
Linux BenchmarkGlobalRead 1.182 ns/op -0.022 ns/op / -1.8% (better)
Linux BenchmarkGlobalWrite 7.798 ns/op -0.033 ns/op / -0.4% (better)
Linux BenchmarkGoroutine 24534 ns/op -4287 ns/op / -14.9% (better)
Linux BenchmarkInterfaceCall 5.879 ns/op +0.053 ns/op / +0.9% (worse)
Linux BenchmarkRuntimeGetG 2.891 ns/op -0.04 ns/op / -1.4% (better)
macOS BenchmarkLookupPCRandom 13.620 ns/op +0.7 ns/op / +5.4% (worse)
macOS BenchmarkMergeCompilerFlags 107 ns/op +2 ns/op / +1.9% (worse)
macOS BenchmarkMergeLinkerFlags 67.660 ns/op 0 ns/op / +0.0%
macOS BenchmarkChannelBuffered 25.810 ns/op -0.84 ns/op / -3.2% (better)
macOS BenchmarkChannelHandoff 6910 ns/op -2323 ns/op / -25.2% (better)
macOS BenchmarkDefer 35.900 ns/op -1.04 ns/op / -2.8% (better)
macOS BenchmarkDirectCall 1.041 ns/op -0.015 ns/op / -1.4% (better)
macOS BenchmarkGlobalRead 1.102 ns/op +0.018 ns/op / +1.7% (worse)
macOS BenchmarkGlobalWrite 1.084 ns/op +0.036 ns/op / +3.4% (worse)
macOS BenchmarkGoroutine 36698 ns/op -2387 ns/op / -6.1% (better)
macOS BenchmarkInterfaceCall 3.988 ns/op +0.072 ns/op / +1.8% (worse)
macOS BenchmarkRuntimeGetG 1.999 ns/op -0.2 ns/op / -9.1% (better)
Windows MinGW BenchmarkLookupPCRandom 13.240 ns/op +0.31 ns/op / +2.4% (worse)
Windows MinGW BenchmarkMergeCompilerFlags 625.200 ns/op -48.1 ns/op / -7.1% (better)
Windows MinGW BenchmarkMergeLinkerFlags 547 ns/op -48.8 ns/op / -8.2% (better)
Windows MinGW BenchmarkChannelBuffered 30.200 ns/op -0.48 ns/op / -1.6% (better)
Windows MinGW BenchmarkChannelHandoff 889.100 ns/op -33 ns/op / -3.6% (better)
Windows MinGW BenchmarkDefer 58.360 ns/op +0.11 ns/op / +0.2% (worse)
Windows MinGW BenchmarkDirectCall 1.548 ns/op -0.001 ns/op / -0.1% (better)
Windows MinGW BenchmarkGlobalRead 1.552 ns/op -0.002 ns/op / -0.1% (better)
Windows MinGW BenchmarkGlobalWrite 2.473 ns/op +0.002 ns/op / +0.1% (worse)
Windows MinGW BenchmarkGoroutine 92769 ns/op +406 ns/op / +0.4% (worse)
Windows MinGW BenchmarkInterfaceCall 8.371 ns/op 0 ns/op / +0.0%
Windows MinGW BenchmarkRuntimeGetG 2.480 ns/op +0.001 ns/op / +0.04034% (worse)
Windows MinGW 386 BenchmarkLookupPCRandom 21.540 ns/op -0.03 ns/op / -0.1% (better)
Windows MinGW 386 BenchmarkMergeCompilerFlags 555.500 ns/op -8.9 ns/op / -1.6% (better)
Windows MinGW 386 BenchmarkMergeLinkerFlags 521.400 ns/op -8 ns/op / -1.5% (better)
Windows MinGW 386 BenchmarkChannelBuffered 33.990 ns/op +0.1 ns/op / +0.3% (worse)
Windows MinGW 386 BenchmarkChannelHandoff 771.700 ns/op -26.5 ns/op / -3.3% (better)
Windows MinGW 386 BenchmarkDefer 35.870 ns/op -0.94 ns/op / -2.6% (better)
Windows MinGW 386 BenchmarkDirectCall 1.362 ns/op +0.003 ns/op / +0.2% (worse)
Windows MinGW 386 BenchmarkGlobalRead 1.365 ns/op +0.005 ns/op / +0.4% (worse)
Windows MinGW 386 BenchmarkGlobalWrite 6.986 ns/op -0.004 ns/op / -0.1% (better)
Windows MinGW 386 BenchmarkGoroutine 69725 ns/op +3 ns/op / +0.004303% (worse)
Windows MinGW 386 BenchmarkInterfaceCall 7.360 ns/op +0.024 ns/op / +0.3% (worse)
Windows MinGW 386 BenchmarkRuntimeGetG 1.637 ns/op +0.005 ns/op / +0.3% (worse)
Windows MinGW ARM64 BenchmarkLookupPCRandom 12.160 ns/op +0.04 ns/op / +0.3% (worse)
Windows MinGW ARM64 BenchmarkMergeCompilerFlags 565 ns/op -2.3 ns/op / -0.4% (better)
Windows MinGW ARM64 BenchmarkMergeLinkerFlags 528.200 ns/op -2 ns/op / -0.4% (better)
Windows MinGW ARM64 BenchmarkChannelBuffered 38.940 ns/op +1.34 ns/op / +3.6% (worse)
Windows MinGW ARM64 BenchmarkChannelHandoff 2133 ns/op -303 ns/op / -12.4% (better)
Windows MinGW ARM64 BenchmarkDefer 52.530 ns/op +3.4 ns/op / +6.9% (worse)
Windows MinGW ARM64 BenchmarkDirectCall 0.664 ns/op 0 ns/op / +0.0%
Windows MinGW ARM64 BenchmarkGlobalRead 0.663 ns/op -0.0001 ns/op / -0.01507% (better)
Windows MinGW ARM64 BenchmarkGlobalWrite 0.737 ns/op +0.0002 ns/op / +0.02715% (worse)
Windows MinGW ARM64 BenchmarkGoroutine 59906 ns/op +1441 ns/op / +2.5% (worse)
Windows MinGW ARM64 BenchmarkInterfaceCall 4.148 ns/op +0.006 ns/op / +0.1% (worse)
Windows MinGW ARM64 BenchmarkRuntimeGetG 1.770 ns/op 0 ns/op / +0.0%
Windows MSVC BenchmarkLookupPCRandom 7.116 ns/op -0.013 ns/op / -0.2% (better)
Windows MSVC BenchmarkMergeCompilerFlags 602.400 ns/op -81.6 ns/op / -11.9% (better)
Windows MSVC BenchmarkMergeLinkerFlags 593.300 ns/op -47 ns/op / -7.3% (better)
Windows MSVC BenchmarkChannelBuffered 28.510 ns/op 0 ns/op / +0.0%
Windows MSVC BenchmarkChannelHandoff 812.500 ns/op +105 ns/op / +14.8% (worse)
Windows MSVC BenchmarkDefer 33.950 ns/op +2.11 ns/op / +6.6% (worse)
Windows MSVC BenchmarkDirectCall 0.910 ns/op +0.0159 ns/op / +1.8% (worse)
Windows MSVC BenchmarkGlobalRead 0.931 ns/op +0.0306 ns/op / +3.4% (worse)
Windows MSVC BenchmarkGlobalWrite 4.475 ns/op +0.007 ns/op / +0.2% (worse)
Windows MSVC BenchmarkGoroutine 40592 ns/op +647 ns/op / +1.6% (worse)
Windows MSVC BenchmarkInterfaceCall 5.105 ns/op +0.149 ns/op / +3.0% (worse)
Windows MSVC BenchmarkRuntimeGetG 1.199 ns/op +0.169 ns/op / +16.4% (worse)
Windows MSVC 386 BenchmarkLookupPCRandom 26.550 ns/op +0.06 ns/op / +0.2% (worse)
Windows MSVC 386 BenchmarkMergeCompilerFlags 743.600 ns/op +37.4 ns/op / +5.3% (worse)
Windows MSVC 386 BenchmarkMergeLinkerFlags 660.100 ns/op -36 ns/op / -5.2% (better)
Windows MSVC 386 BenchmarkChannelBuffered 38.820 ns/op +0.22 ns/op / +0.6% (worse)
Windows MSVC 386 BenchmarkChannelHandoff 847.100 ns/op -41.5 ns/op / -4.7% (better)
Windows MSVC 386 BenchmarkDefer 46.190 ns/op +0.72 ns/op / +1.6% (worse)
Windows MSVC 386 BenchmarkDirectCall 1.548 ns/op +0.001 ns/op / +0.1% (worse)
Windows MSVC 386 BenchmarkGlobalRead 1.552 ns/op +0.002 ns/op / +0.1% (worse)
Windows MSVC 386 BenchmarkGlobalWrite 7.780 ns/op +0.002 ns/op / +0.02571% (worse)
Windows MSVC 386 BenchmarkGoroutine 110737 ns/op -2971 ns/op / -2.6% (better)
Windows MSVC 386 BenchmarkInterfaceCall 8.421 ns/op +0.048 ns/op / +0.6% (worse)
Windows MSVC 386 BenchmarkRuntimeGetG 2.169 ns/op +0.001 ns/op / +0.04613% (worse)
Windows MSVC ARM64 BenchmarkLookupPCRandom 12.020 ns/op -0.15 ns/op / -1.2% (better)
Windows MSVC ARM64 BenchmarkMergeCompilerFlags 568.300 ns/op -2.5 ns/op / -0.4% (better)
Windows MSVC ARM64 BenchmarkMergeLinkerFlags 538.800 ns/op +7.1 ns/op / +1.3% (worse)
Windows MSVC ARM64 BenchmarkChannelBuffered 37.530 ns/op -1.94 ns/op / -4.9% (better)
Windows MSVC ARM64 BenchmarkChannelHandoff 3186 ns/op +99 ns/op / +3.2% (worse)
Windows MSVC ARM64 BenchmarkDefer 63.530 ns/op -0.68 ns/op / -1.1% (better)
Windows MSVC ARM64 BenchmarkDirectCall 0.663 ns/op -0.0001 ns/op / -0.01507% (better)
Windows MSVC ARM64 BenchmarkGlobalRead 0.663 ns/op -0.0001 ns/op / -0.01508% (better)
Windows MSVC ARM64 BenchmarkGlobalWrite 3.796 ns/op -0.001 ns/op / -0.02634% (better)
Windows MSVC ARM64 BenchmarkGoroutine 59807 ns/op -2392 ns/op / -3.8% (better)
Windows MSVC ARM64 BenchmarkInterfaceCall 4.149 ns/op +0.003 ns/op / +0.1% (worse)
Windows MSVC ARM64 BenchmarkRuntimeGetG 1.770 ns/op -0.035 ns/op / -1.9% (better)
Timer runtime benchmarks
Platform Operation and runtime ns/op vs base
Linux AfterFuncZeroDelivery/Go 926.300 ns/op +20.3 ns/op / +2.2% (worse)
Linux AfterFuncZeroDelivery/LLGo 36501 ns/op -1971 ns/op / -5.1% (better)
Linux CreateStop/Go 293.200 ns/op -5.3 ns/op / -1.8% (better)
Linux CreateStop/LLGo 1691 ns/op -204 ns/op / -10.8% (better)
Linux RearmStopped/Go 119.800 ns/op +3.7 ns/op / +3.2% (worse)
Linux RearmStopped/LLGo 1233 ns/op -101 ns/op / -7.6% (better)
Linux ResetActive/Go 68.680 ns/op -0.12 ns/op / -0.2% (better)
Linux ResetActive/LLGo 799.300 ns/op -58.1 ns/op / -6.8% (better)
Linux ResetHeap1024/Go 67.100 ns/op +0.07 ns/op / +0.1% (worse)
Linux ResetHeap1024/LLGo 176.800 ns/op -1.3 ns/op / -0.7% (better)
macOS AfterFuncZeroDelivery/Go 426.200 ns/op -2.8 ns/op / -0.7% (better)
macOS AfterFuncZeroDelivery/LLGo 107138 ns/op +36869 ns/op / +52.5% (worse)
macOS CreateStop/Go 125.600 ns/op -18.1 ns/op / -12.6% (better)
macOS CreateStop/LLGo 487.400 ns/op -37.5 ns/op / -7.1% (better)
macOS RearmStopped/Go 52.900 ns/op -12.78 ns/op / -19.5% (better)
macOS RearmStopped/LLGo 363.400 ns/op +43.8 ns/op / +13.7% (worse)
macOS ResetActive/Go 37.860 ns/op -15.08 ns/op / -28.5% (better)
macOS ResetActive/LLGo 252.100 ns/op +97.5 ns/op / +63.1% (worse)
macOS ResetHeap1024/Go 40.720 ns/op -3.5 ns/op / -7.9% (better)
macOS ResetHeap1024/LLGo 87.770 ns/op +2.65 ns/op / +3.1% (worse)
Windows MinGW AfterFuncZeroDelivery/Go 564.100 ns/op +6.9 ns/op / +1.2% (worse)
Windows MinGW AfterFuncZeroDelivery/LLGo 183596 ns/op -815 ns/op / -0.4% (better)
Windows MinGW CreateStop/Go 115.400 ns/op +0.4 ns/op / +0.3% (worse)
Windows MinGW CreateStop/LLGo 412.500 ns/op -14.3 ns/op / -3.4% (better)
Windows MinGW RearmStopped/Go 31.530 ns/op -0.16 ns/op / -0.5% (better)
Windows MinGW RearmStopped/LLGo 330.700 ns/op +48.9 ns/op / +17.4% (worse)
Windows MinGW ResetActive/Go 20.070 ns/op -0.11 ns/op / -0.5% (better)
Windows MinGW ResetActive/LLGo 171.300 ns/op +7.5 ns/op / +4.6% (worse)
Windows MinGW ResetHeap1024/Go 20.520 ns/op +0.08 ns/op / +0.4% (worse)
Windows MinGW ResetHeap1024/LLGo 126.800 ns/op +3.1 ns/op / +2.5% (worse)
Windows MinGW 386 AfterFuncZeroDelivery/Go 770.600 ns/op +2.9 ns/op / +0.4% (worse)
Windows MinGW 386 AfterFuncZeroDelivery/LLGo 128691 ns/op -416 ns/op / -0.3% (better)
Windows MinGW 386 CreateStop/Go 165.900 ns/op -1.3 ns/op / -0.8% (better)
Windows MinGW 386 CreateStop/LLGo 389.900 ns/op -19.3 ns/op / -4.7% (better)
Windows MinGW 386 RearmStopped/Go 56.720 ns/op -0.18 ns/op / -0.3% (better)
Windows MinGW 386 RearmStopped/LLGo 282.400 ns/op +3.4 ns/op / +1.2% (worse)
Windows MinGW 386 ResetActive/Go 32.560 ns/op -0.02 ns/op / -0.1% (better)
Windows MinGW 386 ResetActive/LLGo 908.300 ns/op +125.8 ns/op / +16.1% (worse)
Windows MinGW 386 ResetHeap1024/Go 32.810 ns/op -0.09 ns/op / -0.3% (better)
Windows MinGW 386 ResetHeap1024/LLGo 150.300 ns/op +1.1 ns/op / +0.7% (worse)
Windows MinGW ARM64 AfterFuncZeroDelivery/Go 663.300 ns/op -4.9 ns/op / -0.7% (better)
Windows MinGW ARM64 AfterFuncZeroDelivery/LLGo 132397 ns/op -3287 ns/op / -2.4% (better)
Windows MinGW ARM64 CreateStop/Go 197.800 ns/op -3.7 ns/op / -1.8% (better)
Windows MinGW ARM64 CreateStop/LLGo 362.500 ns/op +7.2 ns/op / +2.0% (worse)
Windows MinGW ARM64 RearmStopped/Go 70.560 ns/op -0.01 ns/op / -0.01417% (better)
Windows MinGW ARM64 RearmStopped/LLGo 246.400 ns/op -1.8 ns/op / -0.7% (better)
Windows MinGW ARM64 ResetActive/Go 31 ns/op -0.01 ns/op / -0.03225% (better)
Windows MinGW ARM64 ResetActive/LLGo 117.600 ns/op -6.8 ns/op / -5.5% (better)
Windows MinGW ARM64 ResetHeap1024/Go 31.150 ns/op +0.05 ns/op / +0.2% (worse)
Windows MinGW ARM64 ResetHeap1024/LLGo 124.100 ns/op -0.5 ns/op / -0.4% (better)
Windows MSVC AfterFuncZeroDelivery/Go 417.800 ns/op -9.9 ns/op / -2.3% (better)
Windows MSVC AfterFuncZeroDelivery/LLGo 85001 ns/op -194 ns/op / -0.2% (better)
Windows MSVC CreateStop/Go 122 ns/op -1.6 ns/op / -1.3% (better)
Windows MSVC CreateStop/LLGo 308.300 ns/op -0.3 ns/op / -0.1% (better)
Windows MSVC RearmStopped/Go 47.640 ns/op +0.13 ns/op / +0.3% (worse)
Windows MSVC RearmStopped/LLGo 179.600 ns/op -13.3 ns/op / -6.9% (better)
Windows MSVC ResetActive/Go 21.190 ns/op +0.21 ns/op / +1.0% (worse)
Windows MSVC ResetActive/LLGo 510.800 ns/op +127.9 ns/op / +33.4% (worse)
Windows MSVC ResetHeap1024/Go 20.420 ns/op -0.54 ns/op / -2.6% (better)
Windows MSVC ResetHeap1024/LLGo 89.230 ns/op -3.15 ns/op / -3.4% (better)
Windows MSVC 386 AfterFuncZeroDelivery/Go 969.600 ns/op +11.8 ns/op / +1.2% (worse)
Windows MSVC 386 AfterFuncZeroDelivery/LLGo 204158 ns/op +1330 ns/op / +0.7% (worse)
Windows MSVC 386 CreateStop/Go 197 ns/op +4.6 ns/op / +2.4% (worse)
Windows MSVC 386 CreateStop/LLGo 469.300 ns/op +9.9 ns/op / +2.2% (worse)
Windows MSVC 386 RearmStopped/Go 63.490 ns/op +0.1 ns/op / +0.2% (worse)
Windows MSVC 386 RearmStopped/LLGo 318.700 ns/op -2.5 ns/op / -0.8% (better)
Windows MSVC 386 ResetActive/Go 39.030 ns/op +0.01 ns/op / +0.02563% (worse)
Windows MSVC 386 ResetActive/LLGo 948.900 ns/op -14.3 ns/op / -1.5% (better)
Windows MSVC 386 ResetHeap1024/Go 39.480 ns/op +0.09 ns/op / +0.2% (worse)
Windows MSVC 386 ResetHeap1024/LLGo 169.900 ns/op +1.3 ns/op / +0.8% (worse)
Windows MSVC ARM64 AfterFuncZeroDelivery/Go 676.800 ns/op +1.7 ns/op / +0.3% (worse)
Windows MSVC ARM64 AfterFuncZeroDelivery/LLGo 174374 ns/op -2366 ns/op / -1.3% (better)
Windows MSVC ARM64 CreateStop/Go 200.100 ns/op 0 ns/op / +0.0%
Windows MSVC ARM64 CreateStop/LLGo 373.700 ns/op -29.6 ns/op / -7.3% (better)
Windows MSVC ARM64 RearmStopped/Go 70.640 ns/op +0.02 ns/op / +0.02832% (worse)
Windows MSVC ARM64 RearmStopped/LLGo 271.300 ns/op +0.4 ns/op / +0.1% (worse)
Windows MSVC ARM64 ResetActive/Go 31.120 ns/op +0.11 ns/op / +0.4% (worse)
Windows MSVC ARM64 ResetActive/LLGo 129.200 ns/op -3.2 ns/op / -2.4% (better)
Windows MSVC ARM64 ResetHeap1024/Go 31.190 ns/op +0.07 ns/op / +0.2% (worse)
Windows MSVC ARM64 ResetHeap1024/LLGo 136.300 ns/op -1.4 ns/op / -1.0% (better)

Compared with 78973102bf54 measured in the same runner job.

@cpunion
cpunion force-pushed the codex/wasm-debug-runtime-20261006 branch from eed7d78 to ede6941 Compare October 6, 2026 07:58
@cpunion
cpunion force-pushed the codex/wasm-debug-runtime-20261006 branch from ede6941 to 49441f3 Compare October 6, 2026 09:38
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant