Skip to content

feat(wasm): integrate browser filesystem with bounded workers - #2696

Merged
cpunion merged 3 commits into
xgo-dev:mainfrom
cpunion:codex/browser-fs-workers
Sep 30, 2026
Merged

cpunion merged 3 commits into
xgo-dev:mainfrom
cpunion:codex/browser-fs-workers

Conversation

@cpunion

@cpunion cpunion commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator

Browser filesystem calls need to work on every scheduler worker and refer to the same files, descriptors, and cwd as C code. Loading a host shim only in the embedding page leaves worker globals uninitialized or creates independent filesystem state.

Include the browser host in generated Emscripten modules and proxy worker operations to the runtime thread's Module.FS. Allocate response storage on the requesting worker to avoid forbidden Atomics.wait on the browser main thread. Keep Node's native fs, proxy its worker cwd/chdir, and make callback-bridge installation thread-local. Fix fstat to inspect the open node, including after unlink.

Rebased onto main at abe36f6, including merged #2669, #2695, #2539, and #2700; none are outstanding dependencies. Duplicate HTML host and callback bridge changes are supplied by main. Native runtime behavior is unchanged.

Validation on macOS arm64:

  • Memory32/Memory64 × 1/2 workers: Wasm validation, Node execution, and Chrome with COOP/COEP all pass.
  • The fixture asserts use of distinct scheduler workers, shared files/descriptors/cwd, positioned reads, ENOENT, stat after unlink, GC, and browser Go/C interoperability.
  • Existing worker scheduler/hardening tests and both public llgo test worker paths pass; TestPoolAfterGC passes.
  • Host-shim and Emscripten integration unit tests pass.

The four filesystem configurations now run in the existing worker CI job. The proxy transfers byte payloads through bounded shared memory and serializes only metadata. Filesystem persistence and throughput tuning remain separate work.

Review follow-up: allow-list host methods, install worker wrappers once, validate buffer ranges and returned byte counts, preserve partial-read offsets, and transfer read/write bytes through shared memory instead of JSON arrays. Memory32/64 proxy regressions cover malformed requests, callbacks, cleanup and repeated installation. All four Node/Chrome filesystem configurations pass again after rebase. The collector segment lookup/sweep improvement now comes from main via #2695; its duplicate contribution commit has been removed.

Memory-growth follow-up: refresh the runtime thread's memory views before decoding a worker request and checking its shared payload. The proxy regression now uses shared WebAssembly.Memory and forces worker-side allocation to grow memory for both large writes and reads, with offsets, byte contents, surrounding bytes, cleanup, and metadata-only JSON checked in both pointer widths. The new regression reproduces EINVAL before the fix and passes afterward. All four filesystem configurations pass Wasm validation, Node, and Chrome on this rebased head.

@fennoai fennoai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review: share browser filesystem across scheduler workers

Solid, well-tested change. The build-tag matrix (_default/_wasi_threads/_stub pairs) is disciplined, the segmented-heap refactor threads segment.first-relative arithmetic consistently, and the STW protocol in wasi_gc_world.c is carefully guarded (excludes the requesting thread from the wait count, timeout fallback). No correctness-blocking defects found. A few items worth considering, inline and below.

Performance — segmented heap scan on the targeted path. segmentForBlock/segmentForAddress (segments.go:49, :60) are O(segments) linear scans, and they're called from the innermost GC primitives (gcStateOf, gcAddressOf) for every block in sweep/finishMark and inside Alloc's scan loop. Several call sites re-resolve the segment 2–3× for the same block (gcSetState, gcMarkFree). This PR's whole premise is that workers blocked in host FS calls force growHeap to add disjoint arenas (up to maxHeapSegments = 128), so under the exact workload it targets, per-block ops degrade toward O(blocks × segments). Consider hoisting the segment pointer out of the per-block loops (it's loop-invariant within sweep/finishMark) and passing it into the helpers rather than re-resolving.

Performance — FS proxy serializes byte payloads through JSON. Every proxied read/write embeds bytes as a decimal JSON array via Array.from (browser_fs.js:9, :97) plus multiple copies, on a synchronous cross-thread round-trip. Bulk I/O will be dominated by JSON encoding rather than the FS work. Consider transferring read/write payloads through shared linear memory (offset/length) and reserving JSON for the small metadata envelope.

Clarity — gc_wasm.c arena reservation. llgo_wasi_gc_init_arena() mallocs a fixed 32 MiB region backing llgo_gc_heap_base()/llgo_gc_memory_size(), but llgo_gc_new_arena(size) just malloc(size) and ignores those globals — so the _start/_end globals only ever describe segment 0 and llgo_gc_grow_memory can never actually grow on the wasi-threads path. The naming/comments imply a reserved region that later arenas carve from; either simplify to a per-segment malloc model or document that growth is segment-driven.

Comment thread runtime/internal/runtime/tinygogc/segments.go Outdated
Comment thread internal/wasmworkers/browser_fs.js Outdated
Comment thread internal/wasmworkers/browser_fs.js
Comment thread internal/wasmworkers/browser_fs.js
@codecov

codecov Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@github-actions

github-actions Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

LLGo baseline benchmarks

c627831ed7db | workflow run | long-term charts

Program measurements

Platform Workload File size vs base Text size vs base Build vs base Run vs base
Linux cprintf 7160 B 0 B / +0.0% 387 B 0 B / +0.0% 563.198 ms -1.128 ms / -0.2% (better) 1.373 ms +153.5 us / +12.6% (worse)
Linux cprintf-lto 6912 B 0 B / +0.0% 368 B 0 B / +0.0% 559.996 ms -49.43 ms / -8.1% (better) 1.278 ms +45.59 us / +3.7% (worse)
Linux fmtprintf 1663312 B +1040 B / +0.1% (worse) 498386 B +27 B / +0.005418% (worse) 4.117 s -79.84 ms / -1.9% (better) 3.190 ms +47.13 us / +1.5% (worse)
Linux fmtprintf-lto 1501104 B +720 B / +0.04799% (worse) 436123 B +10 B / +0.002293% (worse) 11.060 s -298.8 ms / -2.6% (better) 2.885 ms -38.15 us / -1.3% (better)
Linux println 69232 B +240 B / +0.3% (worse) 16805 B +22 B / +0.1% (worse) 571.551 ms -23.71 ms / -4.0% (better) 1.609 ms +29.95 us / +1.9% (worse)
Linux println-lto 59960 B +112 B / +0.2% (worse) 14209 B +10 B / +0.1% (worse) 855.309 ms -2.742 ms / -0.3% (better) 1.624 ms +49.86 us / +3.2% (worse)
macOS cprintf 68064 B 0 B / +0.0% 4429 B 0 B / +0.0% 1.014 s -90.47 ms / -8.2% (better) 3.444 ms -3.129 ms / -47.6% (better)
macOS cprintf-lto 84288 B 0 B / +0.0% 193 B 0 B / +0.0% 1.131 s -16.3 ms / -1.4% (better) 3.858 ms -4.859 ms / -55.7% (better)
macOS fmtprintf 1504832 B +160 B / +0.01063% (worse) 875320 B +748 B / +0.1% (worse) 4.066 s -3.504 s / -46.3% (better) 8.705 ms -4.372 ms / -33.4% (better)
macOS fmtprintf-lto 1192704 B 0 B / +0.0% 848952 B +736 B / +0.1% (worse) 10.215 s -167.9 ms / -1.6% (better) 6.495 ms -2.018 ms / -23.7% (better)
macOS println 117216 B +80 B / +0.1% (worse) 37549 B +128 B / +0.3% (worse) 900.050 ms -396.4 ms / -30.6% (better) 5.175 ms -3.149 ms / -37.8% (better)
macOS println-lto 119472 B 0 B / +0.0% 34928 B +104 B / +0.3% (worse) 1.196 s -605.2 ms / -33.6% (better) 4.519 ms +23.33 us / +0.5% (worse)
Windows MinGW cprintf 19456 B 0 B / +0.0% 4550 B 0 B / +0.0% 1.371 s +45.34 ms / +3.4% (worse) 3.676 ms -9.6 us / -0.3% (better)
Windows MinGW cprintf-lto 17920 B 0 B / +0.0% 4486 B 0 B / +0.0% 1.424 s +35.61 ms / +2.6% (worse) 4.018 ms -271.3 us / -6.3% (better)
Windows MinGW fmtprintf 1933312 B +512 B / +0.02649% (worse) 598950 B 0 B / +0.0% 4.491 s +79.28 ms / +1.8% (worse) 8.578 ms -121.5 us / -1.4% (better)
Windows MinGW fmtprintf-lto 1957888 B +512 B / +0.02616% (worse) 547318 B 0 B / +0.0% 10.708 s +210.1 ms / +2.0% (worse) 9.142 ms +89.5 us / +1.0% (worse)
Windows MinGW println 75776 B 0 B / +0.0% 25142 B 0 B / +0.0% 1.369 s +23.78 ms / +1.8% (worse) 7.621 ms +495.5 us / +7.0% (worse)
Windows MinGW println-lto 69120 B 0 B / +0.0% 21990 B 0 B / +0.0% 1.614 s +45.01 ms / +2.9% (worse) 7.681 ms +739.5 us / +10.7% (worse)
Windows MinGW 386 cprintf 43520 B 0 B / +0.0% 5326 B 0 B / +0.0% 1.253 s -52.77 ms / -4.0% (better) 4.969 ms -851.1 us / -14.6% (better)
Windows MinGW 386 cprintf-lto 20992 B 0 B / +0.0% 5094 B 0 B / +0.0% 1.303 s -24.16 ms / -1.8% (better) 4.972 ms -200 ns / -0.004022% (better)
Windows MinGW 386 fmtprintf 1896448 B 0 B / +0.0% 472430 B +16 B / +0.003387% (worse) 4.410 s -29.4 ms / -0.7% (better) 11.441 ms -78 us / -0.7% (better)
Windows MinGW 386 fmtprintf-lto 2181120 B +512 B / +0.02348% (worse) 451274 B +16 B / +0.003546% (worse) 9.934 s -222.2 ms / -2.2% (better) 10.271 ms -42.5 us / -0.4% (better)
Windows MinGW 386 println 96256 B 0 B / +0.0% 21474 B +16 B / +0.1% (worse) 1.357 s -38.06 ms / -2.7% (better) 8.701 ms -670.2 us / -7.2% (better)
Windows MinGW 386 println-lto 74240 B 0 B / +0.0% 19334 B +20 B / +0.1% (worse) 1.491 s -152.9 ms / -9.3% (better) 8.480 ms -2.108 ms / -19.9% (better)
Windows MinGW ARM64 cprintf 18944 B 0 B / +0.0% 4408 B 0 B / +0.0% 1.471 s -8.031 ms / -0.5% (better) 6.024 ms -142.7 us / -2.3% (better)
Windows MinGW ARM64 cprintf-lto 17920 B 0 B / +0.0% 4340 B 0 B / +0.0% 1.505 s -26.41 ms / -1.7% (better) 6.024 ms +129 us / +2.2% (worse)
Windows MinGW ARM64 fmtprintf 1819648 B +512 B / +0.02815% (worse) 509616 B +8 B / +0.00157% (worse) 4.278 s +70 ms / +1.7% (worse) 11.741 ms -5.1 us / -0.04342% (better)
Windows MinGW ARM64 fmtprintf-lto 1880064 B +1024 B / +0.1% (worse) 476176 B +24 B / +0.00504% (worse) 9.568 s +172.3 ms / +1.8% (worse) 12.760 ms +795.5 us / +6.6% (worse)
Windows MinGW ARM64 println 72192 B 0 B / +0.0% 23884 B +8 B / +0.03351% (worse) 1.499 s +30.11 ms / +2.1% (worse) 10.676 ms +256.6 us / +2.5% (worse)
Windows MinGW ARM64 println-lto 68608 B 0 B / +0.0% 21232 B +8 B / +0.03769% (worse) 1.686 s -18.99 ms / -1.1% (better) 10.514 ms +352.3 us / +3.5% (worse)
Windows MSVC cprintf 120320 B 0 B / +0.0% 65798 B 0 B / +0.0% 1.186 s +86.45 ms / +7.9% (worse) 4.852 ms +1.435 ms / +42.0% (worse)
Windows MSVC cprintf-lto 119808 B 0 B / +0.0% 65734 B 0 B / +0.0% 1.135 s -171.5 ms / -13.1% (better) 3.321 ms -57.8 us / -1.7% (better)
Windows MSVC fmtprintf 1643520 B +512 B / +0.03116% (worse) 694502 B 0 B / +0.0% 4.157 s -198.2 ms / -4.6% (better) 10.259 ms -516.6 us / -4.8% (better)
Windows MSVC fmtprintf-lto 1635328 B +512 B / +0.03132% (worse) 647078 B +16 B / +0.002473% (worse) 9.583 s -686.2 ms / -6.7% (better) 11.111 ms +299.2 us / +2.8% (worse)
Windows MSVC println 194560 B 0 B / +0.0% 120822 B 0 B / +0.0% 1.122 s -10.37 ms / -0.9% (better) 7.845 ms -118.3 us / -1.5% (better)
Windows MSVC println-lto 192512 B 0 B / +0.0% 118358 B +16 B / +0.01352% (worse) 1.320 s -64.24 ms / -4.6% (better) 7.441 ms -704.6 us / -8.7% (better)
Windows MSVC 386 cprintf 9728 B 0 B / +0.0% 3931 B 0 B / +0.0% 851.488 ms +6.092 ms / +0.7% (worse) 4.317 ms +10.3 us / +0.2% (worse)
Windows MSVC 386 cprintf-lto 9216 B 0 B / +0.0% 3853 B 0 B / +0.0% 903.261 ms +42.59 ms / +4.9% (worse) 4.381 ms -2.218 ms / -33.6% (better)
Windows MSVC 386 fmtprintf 1204736 B +512 B / +0.04252% (worse) 455820 B +16 B / +0.00351% (worse) 3.235 s -15.2 ms / -0.5% (better) 8.619 ms -438.2 us / -4.8% (better)
Windows MSVC 386 fmtprintf-lto 1241600 B +512 B / +0.04125% (worse) 427211 B +16 B / +0.003745% (worse) 7.318 s +15.25 ms / +0.2% (worse) 8.886 ms -535.6 us / -5.7% (better)
Windows MSVC 386 println 36352 B 0 B / +0.0% 20324 B 0 B / +0.0% 901.430 ms +36.86 ms / +4.3% (worse) 7.378 ms +197.1 us / +2.7% (worse)
Windows MSVC 386 println-lto 35840 B 0 B / +0.0% 18565 B +16 B / +0.1% (worse) 1.087 s +73.35 ms / +7.2% (worse) 7.389 ms -179.5 us / -2.4% (better)
Windows MSVC ARM64 cprintf 11776 B 0 B / +0.0% 4192 B 0 B / +0.0% 1.159 s -8.748 ms / -0.7% (better) 6.465 ms +224.9 us / +3.6% (worse)
Windows MSVC ARM64 cprintf-lto 10752 B 0 B / +0.0% 4084 B 0 B / +0.0% 1.178 s +17.51 ms / +1.5% (worse) 6.295 ms +199.4 us / +3.3% (worse)
Windows MSVC ARM64 fmtprintf 1387008 B +512 B / +0.03693% (worse) 509560 B +16 B / +0.00314% (worse) 3.811 s +46.79 ms / +1.2% (worse) 13.525 ms -65.8 us / -0.5% (better)
Windows MSVC ARM64 fmtprintf-lto 1405440 B +512 B / +0.03644% (worse) 476836 B +16 B / +0.003356% (worse) 8.508 s -7.993 ms / -0.1% (better) 13.068 ms -128.6 us / -1.0% (better)
Windows MSVC ARM64 println 45056 B 0 B / +0.0% 23908 B 0 B / +0.0% 1.179 s +32.38 ms / +2.8% (worse) 11.435 ms +421.5 us / +3.8% (worse)
Windows MSVC ARM64 println-lto 42496 B 0 B / +0.0% 21380 B 0 B / +0.0% 1.384 s +47.36 ms / +3.5% (worse) 11.929 ms +173 us / +1.5% (worse)
Core language and compiler benchmarks
Platform Benchmark ns/op vs base
Linux BenchmarkLookupPCRandom 14.550 ns/op -0.2 ns/op / -1.4% (better)
Linux BenchmarkMergeCompilerFlags 197.200 ns/op +0.4 ns/op / +0.2% (worse)
Linux BenchmarkMergeLinkerFlags 135.600 ns/op +4.6 ns/op / +3.5% (worse)
Linux BenchmarkChannelBuffered 55.500 ns/op -0.07 ns/op / -0.1% (better)
Linux BenchmarkChannelHandoff 19321 ns/op +6005 ns/op / +45.1% (worse)
Linux BenchmarkDefer 47.840 ns/op -3.67 ns/op / -7.1% (better)
Linux BenchmarkDirectCall 1.598 ns/op +0.043 ns/op / +2.8% (worse)
Linux BenchmarkGlobalRead 1.169 ns/op +0.005 ns/op / +0.4% (worse)
Linux BenchmarkGlobalWrite 7.759 ns/op -0.003 ns/op / -0.03865% (better)
Linux BenchmarkGoroutine 24903 ns/op +692 ns/op / +2.9% (worse)
Linux BenchmarkInterfaceCall 6.067 ns/op +0.195 ns/op / +3.3% (worse)
Linux BenchmarkRuntimeGetG 2.893 ns/op -0.116 ns/op / -3.9% (better)
macOS BenchmarkLookupPCRandom 14.900 ns/op +0.36 ns/op / +2.5% (worse)
macOS BenchmarkMergeCompilerFlags 140.600 ns/op -2.6 ns/op / -1.8% (better)
macOS BenchmarkMergeLinkerFlags 86.860 ns/op -0.08 ns/op / -0.1% (better)
macOS BenchmarkChannelBuffered 30.370 ns/op -6.39 ns/op / -17.4% (better)
macOS BenchmarkChannelHandoff 10968 ns/op -945 ns/op / -7.9% (better)
macOS BenchmarkDefer 44.120 ns/op -4.29 ns/op / -8.9% (better)
macOS BenchmarkDirectCall 1.306 ns/op +0.089 ns/op / +7.3% (worse)
macOS BenchmarkGlobalRead 1.302 ns/op -0.071 ns/op / -5.2% (better)
macOS BenchmarkGlobalWrite 1.482 ns/op +0.204 ns/op / +16.0% (worse)
macOS BenchmarkGoroutine 95991 ns/op +33419 ns/op / +53.4% (worse)
macOS BenchmarkInterfaceCall 4.672 ns/op -0.808 ns/op / -14.7% (better)
macOS BenchmarkRuntimeGetG 2.513 ns/op -1.106 ns/op / -30.6% (better)
Windows MinGW BenchmarkLookupPCRandom 12.980 ns/op -0.17 ns/op / -1.3% (better)
Windows MinGW BenchmarkMergeCompilerFlags 619 ns/op -7.7 ns/op / -1.2% (better)
Windows MinGW BenchmarkMergeLinkerFlags 556.900 ns/op +6.6 ns/op / +1.2% (worse)
Windows MinGW BenchmarkChannelBuffered 30.130 ns/op +0.22 ns/op / +0.7% (worse)
Windows MinGW BenchmarkChannelHandoff 926.700 ns/op +58 ns/op / +6.7% (worse)
Windows MinGW BenchmarkDefer 60.210 ns/op +0.8 ns/op / +1.3% (worse)
Windows MinGW BenchmarkDirectCall 1.547 ns/op -0.002 ns/op / -0.1% (better)
Windows MinGW BenchmarkGlobalRead 1.858 ns/op -0.004 ns/op / -0.2% (better)
Windows MinGW BenchmarkGlobalWrite 2.459 ns/op -0.003 ns/op / -0.1% (better)
Windows MinGW BenchmarkGoroutine 103184 ns/op +12281 ns/op / +13.5% (worse)
Windows MinGW BenchmarkInterfaceCall 8.386 ns/op +0.006 ns/op / +0.1% (worse)
Windows MinGW BenchmarkRuntimeGetG 2.171 ns/op +0.004 ns/op / +0.2% (worse)
Windows MinGW 386 BenchmarkLookupPCRandom 26.600 ns/op +0.04 ns/op / +0.2% (worse)
Windows MinGW 386 BenchmarkMergeCompilerFlags 769.300 ns/op +34.8 ns/op / +4.7% (worse)
Windows MinGW 386 BenchmarkMergeLinkerFlags 706.200 ns/op +50.4 ns/op / +7.7% (worse)
Windows MinGW 386 BenchmarkChannelBuffered 38.740 ns/op -0.39 ns/op / -1.0% (better)
Windows MinGW 386 BenchmarkChannelHandoff 909.500 ns/op +53.5 ns/op / +6.3% (worse)
Windows MinGW 386 BenchmarkDefer 43.700 ns/op +0.84 ns/op / +2.0% (worse)
Windows MinGW 386 BenchmarkDirectCall 1.548 ns/op -0.001 ns/op / -0.1% (better)
Windows MinGW 386 BenchmarkGlobalRead 1.548 ns/op 0 ns/op / +0.0%
Windows MinGW 386 BenchmarkGlobalWrite 7.779 ns/op +0.01 ns/op / +0.1% (worse)
Windows MinGW 386 BenchmarkGoroutine 104752 ns/op -2067 ns/op / -1.9% (better)
Windows MinGW 386 BenchmarkInterfaceCall 8.386 ns/op +0.006 ns/op / +0.1% (worse)
Windows MinGW 386 BenchmarkRuntimeGetG 2.166 ns/op -0.003 ns/op / -0.1% (better)
Windows MinGW ARM64 BenchmarkLookupPCRandom 12.070 ns/op +0.01 ns/op / +0.1% (worse)
Windows MinGW ARM64 BenchmarkMergeCompilerFlags 569.500 ns/op -4 ns/op / -0.7% (better)
Windows MinGW ARM64 BenchmarkMergeLinkerFlags 534.900 ns/op -3.8 ns/op / -0.7% (better)
Windows MinGW ARM64 BenchmarkChannelBuffered 37.830 ns/op -1.44 ns/op / -3.7% (better)
Windows MinGW ARM64 BenchmarkChannelHandoff 1921 ns/op -24 ns/op / -1.2% (better)
Windows MinGW ARM64 BenchmarkDefer 53.430 ns/op +0.47 ns/op / +0.9% (worse)
Windows MinGW ARM64 BenchmarkDirectCall 0.589 ns/op -0.0011 ns/op / -0.2% (better)
Windows MinGW ARM64 BenchmarkGlobalRead 0.663 ns/op -0.0003 ns/op / -0.04521% (better)
Windows MinGW ARM64 BenchmarkGlobalWrite 0.663 ns/op -0.0004 ns/op / -0.1% (better)
Windows MinGW ARM64 BenchmarkGoroutine 60661 ns/op +918 ns/op / +1.5% (worse)
Windows MinGW ARM64 BenchmarkInterfaceCall 4.156 ns/op +0.012 ns/op / +0.3% (worse)
Windows MinGW ARM64 BenchmarkRuntimeGetG 1.778 ns/op +0.008 ns/op / +0.5% (worse)
Windows MSVC BenchmarkLookupPCRandom 13.120 ns/op -0.13 ns/op / -1.0% (better)
Windows MSVC BenchmarkMergeCompilerFlags 651.200 ns/op -14.8 ns/op / -2.2% (better)
Windows MSVC BenchmarkMergeLinkerFlags 545.100 ns/op -68 ns/op / -11.1% (better)
Windows MSVC BenchmarkChannelBuffered 30.770 ns/op -1.18 ns/op / -3.7% (better)
Windows MSVC BenchmarkChannelHandoff 1160 ns/op -16 ns/op / -1.4% (better)
Windows MSVC BenchmarkDefer 54.470 ns/op -4.26 ns/op / -7.3% (better)
Windows MSVC BenchmarkDirectCall 1.548 ns/op -0.004 ns/op / -0.3% (better)
Windows MSVC BenchmarkGlobalRead 1.548 ns/op -0.004 ns/op / -0.3% (better)
Windows MSVC BenchmarkGlobalWrite 2.467 ns/op -0.006 ns/op / -0.2% (better)
Windows MSVC BenchmarkGoroutine 86080 ns/op -3165 ns/op / -3.5% (better)
Windows MSVC BenchmarkInterfaceCall 8.392 ns/op +0.023 ns/op / +0.3% (worse)
Windows MSVC BenchmarkRuntimeGetG 1.863 ns/op +0.001 ns/op / +0.1% (worse)
Windows MSVC 386 BenchmarkLookupPCRandom 21.550 ns/op -0.01 ns/op / -0.04638% (better)
Windows MSVC 386 BenchmarkMergeCompilerFlags 548.800 ns/op -5.1 ns/op / -0.9% (better)
Windows MSVC 386 BenchmarkMergeLinkerFlags 514.800 ns/op -3 ns/op / -0.6% (better)
Windows MSVC 386 BenchmarkChannelBuffered 33.880 ns/op -0.08 ns/op / -0.2% (better)
Windows MSVC 386 BenchmarkChannelHandoff 686.700 ns/op +76.2 ns/op / +12.5% (worse)
Windows MSVC 386 BenchmarkDefer 38.950 ns/op +0.04 ns/op / +0.1% (worse)
Windows MSVC 386 BenchmarkDirectCall 1.358 ns/op +0.002 ns/op / +0.1% (worse)
Windows MSVC 386 BenchmarkGlobalRead 1.358 ns/op -0.001 ns/op / -0.1% (better)
Windows MSVC 386 BenchmarkGlobalWrite 6.979 ns/op -0.003 ns/op / -0.04297% (better)
Windows MSVC 386 BenchmarkGoroutine 73240 ns/op +829 ns/op / +1.1% (worse)
Windows MSVC 386 BenchmarkInterfaceCall 7.329 ns/op -0.012 ns/op / -0.2% (better)
Windows MSVC 386 BenchmarkRuntimeGetG 1.901 ns/op -0.006 ns/op / -0.3% (better)
Windows MSVC ARM64 BenchmarkLookupPCRandom 12.080 ns/op +0.04 ns/op / +0.3% (worse)
Windows MSVC ARM64 BenchmarkMergeCompilerFlags 573.600 ns/op +7 ns/op / +1.2% (worse)
Windows MSVC ARM64 BenchmarkMergeLinkerFlags 535.900 ns/op +4.1 ns/op / +0.8% (worse)
Windows MSVC ARM64 BenchmarkChannelBuffered 37.640 ns/op -1.26 ns/op / -3.2% (better)
Windows MSVC ARM64 BenchmarkChannelHandoff 3090 ns/op +373 ns/op / +13.7% (worse)
Windows MSVC ARM64 BenchmarkDefer 59.330 ns/op -0.43 ns/op / -0.7% (better)
Windows MSVC ARM64 BenchmarkDirectCall 0.589 ns/op -0.0006 ns/op / -0.1% (better)
Windows MSVC ARM64 BenchmarkGlobalRead 0.663 ns/op -0.0006 ns/op / -0.1% (better)
Windows MSVC ARM64 BenchmarkGlobalWrite 3.749 ns/op -0.007 ns/op / -0.2% (better)
Windows MSVC ARM64 BenchmarkGoroutine 57388 ns/op -780 ns/op / -1.3% (better)
Windows MSVC ARM64 BenchmarkInterfaceCall 4.137 ns/op -0.017 ns/op / -0.4% (better)
Windows MSVC ARM64 BenchmarkRuntimeGetG 1.802 ns/op +0.033 ns/op / +1.9% (worse)
Timer runtime benchmarks
Platform Operation and runtime ns/op vs base
Linux AfterFuncZeroDelivery/Go 911.500 ns/op -4 ns/op / -0.4% (better)
Linux AfterFuncZeroDelivery/LLGo 37826 ns/op -3981 ns/op / -9.5% (better)
Linux CreateStop/Go 291.600 ns/op +1.5 ns/op / +0.5% (worse)
Linux CreateStop/LLGo 1722 ns/op -54 ns/op / -3.0% (better)
Linux RearmStopped/Go 115 ns/op -1.1 ns/op / -0.9% (better)
Linux RearmStopped/LLGo 1061 ns/op -294 ns/op / -21.7% (better)
Linux ResetActive/Go 67.460 ns/op -1.47 ns/op / -2.1% (better)
Linux ResetActive/LLGo 782.400 ns/op +36.1 ns/op / +4.8% (worse)
Linux ResetHeap1024/Go 67.300 ns/op +0.15 ns/op / +0.2% (worse)
Linux ResetHeap1024/LLGo 178.200 ns/op +4.1 ns/op / +2.4% (worse)
macOS AfterFuncZeroDelivery/Go 550.300 ns/op +48 ns/op / +9.6% (worse)
macOS AfterFuncZeroDelivery/LLGo 95621 ns/op -1857 ns/op / -1.9% (better)
macOS CreateStop/Go 162.200 ns/op +13.7 ns/op / +9.2% (worse)
macOS CreateStop/LLGo 549.600 ns/op -5.5 ns/op / -1.0% (better)
macOS RearmStopped/Go 67.640 ns/op +2.93 ns/op / +4.5% (worse)
macOS RearmStopped/LLGo 338.900 ns/op -77.6 ns/op / -18.6% (better)
macOS ResetActive/Go 47.520 ns/op +1.25 ns/op / +2.7% (worse)
macOS ResetActive/LLGo 165.900 ns/op -90.3 ns/op / -35.2% (better)
macOS ResetHeap1024/Go 50.860 ns/op +4.36 ns/op / +9.4% (worse)
macOS ResetHeap1024/LLGo 94.920 ns/op +1.24 ns/op / +1.3% (worse)
Windows MinGW AfterFuncZeroDelivery/Go 580.300 ns/op +20.8 ns/op / +3.7% (worse)
Windows MinGW AfterFuncZeroDelivery/LLGo 183910 ns/op +1313 ns/op / +0.7% (worse)
Windows MinGW CreateStop/Go 119.600 ns/op +5.6 ns/op / +4.9% (worse)
Windows MinGW CreateStop/LLGo 457.800 ns/op +5.9 ns/op / +1.3% (worse)
Windows MinGW RearmStopped/Go 31.430 ns/op -0.33 ns/op / -1.0% (better)
Windows MinGW RearmStopped/LLGo 282.900 ns/op +2.9 ns/op / +1.0% (worse)
Windows MinGW ResetActive/Go 20.060 ns/op -0.16 ns/op / -0.8% (better)
Windows MinGW ResetActive/LLGo 160.800 ns/op -3.1 ns/op / -1.9% (better)
Windows MinGW ResetHeap1024/Go 20.440 ns/op -0.02 ns/op / -0.1% (better)
Windows MinGW ResetHeap1024/LLGo 124.400 ns/op +0.6 ns/op / +0.5% (worse)
Windows MinGW 386 AfterFuncZeroDelivery/Go 972.800 ns/op +9.6 ns/op / +1.0% (worse)
Windows MinGW 386 AfterFuncZeroDelivery/LLGo 200291 ns/op +1951 ns/op / +1.0% (worse)
Windows MinGW 386 CreateStop/Go 192.500 ns/op +0.7 ns/op / +0.4% (worse)
Windows MinGW 386 CreateStop/LLGo 500.200 ns/op -29.3 ns/op / -5.5% (better)
Windows MinGW 386 RearmStopped/Go 63.530 ns/op -0.24 ns/op / -0.4% (better)
Windows MinGW 386 RearmStopped/LLGo 343.600 ns/op -11.4 ns/op / -3.2% (better)
Windows MinGW 386 ResetActive/Go 39.090 ns/op -0.05 ns/op / -0.1% (better)
Windows MinGW 386 ResetActive/LLGo 993 ns/op +72.2 ns/op / +7.8% (worse)
Windows MinGW 386 ResetHeap1024/Go 39.510 ns/op -0.29 ns/op / -0.7% (better)
Windows MinGW 386 ResetHeap1024/LLGo 187.300 ns/op +0.2 ns/op / +0.1% (worse)
Windows MinGW ARM64 AfterFuncZeroDelivery/Go 668.400 ns/op +3.1 ns/op / +0.5% (worse)
Windows MinGW ARM64 AfterFuncZeroDelivery/LLGo 145218 ns/op +2797 ns/op / +2.0% (worse)
Windows MinGW ARM64 CreateStop/Go 197.100 ns/op -2 ns/op / -1.0% (better)
Windows MinGW ARM64 CreateStop/LLGo 374.400 ns/op +1.5 ns/op / +0.4% (worse)
Windows MinGW ARM64 RearmStopped/Go 70.580 ns/op -0.02 ns/op / -0.02833% (better)
Windows MinGW ARM64 RearmStopped/LLGo 254.900 ns/op +2.8 ns/op / +1.1% (worse)
Windows MinGW ARM64 ResetActive/Go 30.830 ns/op -0.14 ns/op / -0.5% (better)
Windows MinGW ARM64 ResetActive/LLGo 130 ns/op +6.2 ns/op / +5.0% (worse)
Windows MinGW ARM64 ResetHeap1024/Go 31.100 ns/op +0.1 ns/op / +0.3% (worse)
Windows MinGW ARM64 ResetHeap1024/LLGo 127.400 ns/op +0.8 ns/op / +0.6% (worse)
Windows MSVC AfterFuncZeroDelivery/Go 545.400 ns/op -13.5 ns/op / -2.4% (better)
Windows MSVC AfterFuncZeroDelivery/LLGo 167034 ns/op -291 ns/op / -0.2% (better)
Windows MSVC CreateStop/Go 115.800 ns/op -4.2 ns/op / -3.5% (better)
Windows MSVC CreateStop/LLGo 404.600 ns/op -63.1 ns/op / -13.5% (better)
Windows MSVC RearmStopped/Go 31.310 ns/op -0.1 ns/op / -0.3% (better)
Windows MSVC RearmStopped/LLGo 250.200 ns/op -5.6 ns/op / -2.2% (better)
Windows MSVC ResetActive/Go 20.050 ns/op 0 ns/op / +0.0%
Windows MSVC ResetActive/LLGo 142.700 ns/op -12.4 ns/op / -8.0% (better)
Windows MSVC ResetHeap1024/Go 20.430 ns/op -0.17 ns/op / -0.8% (better)
Windows MSVC ResetHeap1024/LLGo 126.800 ns/op -2.2 ns/op / -1.7% (better)
Windows MSVC 386 AfterFuncZeroDelivery/Go 793.400 ns/op +31.2 ns/op / +4.1% (worse)
Windows MSVC 386 AfterFuncZeroDelivery/LLGo 131671 ns/op +1241 ns/op / +1.0% (worse)
Windows MSVC 386 CreateStop/Go 169.200 ns/op +2.8 ns/op / +1.7% (worse)
Windows MSVC 386 CreateStop/LLGo 366.100 ns/op +7.1 ns/op / +2.0% (worse)
Windows MSVC 386 RearmStopped/Go 56.780 ns/op -1.57 ns/op / -2.7% (better)
Windows MSVC 386 RearmStopped/LLGo 261.900 ns/op +1.1 ns/op / +0.4% (worse)
Windows MSVC 386 ResetActive/Go 32.620 ns/op +0.01 ns/op / +0.03067% (worse)
Windows MSVC 386 ResetActive/LLGo 882.500 ns/op +3.3 ns/op / +0.4% (worse)
Windows MSVC 386 ResetHeap1024/Go 32.740 ns/op -0.1 ns/op / -0.3% (better)
Windows MSVC 386 ResetHeap1024/LLGo 142.700 ns/op 0 ns/op / +0.0%
Windows MSVC ARM64 AfterFuncZeroDelivery/Go 674.800 ns/op +4.7 ns/op / +0.7% (worse)
Windows MSVC ARM64 AfterFuncZeroDelivery/LLGo 136981 ns/op -44 ns/op / -0.03211% (better)
Windows MSVC ARM64 CreateStop/Go 207.500 ns/op +10.1 ns/op / +5.1% (worse)
Windows MSVC ARM64 CreateStop/LLGo 376.100 ns/op -4.7 ns/op / -1.2% (better)
Windows MSVC ARM64 RearmStopped/Go 70.570 ns/op -0.09 ns/op / -0.1% (better)
Windows MSVC ARM64 RearmStopped/LLGo 268.200 ns/op +1.4 ns/op / +0.5% (worse)
Windows MSVC ARM64 ResetActive/Go 30.970 ns/op +0.09 ns/op / +0.3% (worse)
Windows MSVC ARM64 ResetActive/LLGo 132.100 ns/op -4.2 ns/op / -3.1% (better)
Windows MSVC ARM64 ResetHeap1024/Go 31.050 ns/op 0 ns/op / +0.0%
Windows MSVC ARM64 ResetHeap1024/LLGo 136.800 ns/op -0.4 ns/op / -0.3% (better)

Compared with 9853dc8d3a4d measured in the same runner job.

@cpunion
cpunion force-pushed the codex/browser-fs-workers branch from 0d9e322 to 14c1d0c Compare September 29, 2026 04:52
@cpunion

cpunion commented Sep 29, 2026

Copy link
Copy Markdown
Collaborator Author

Addressed all four inline findings and the summary notes; rebased onto current main after #2539 merged. The remaining dependency is #2669 and this PR is ready for review.

  • Reads validate caller offsets/lengths and host byte counts and preserve untouched buffer bytes. A malformed result becomes EIO.
  • Read/write bytes now use bounded shared-memory payloads; JSON carries metadata only.
  • Dispatch uses an explicit method allow-list and worker installation is idempotent.
  • The collector segment optimization is included here as well; the first-arena bounds are documented as segment 0 rather than a reservation for later arenas.

Proxy contract tests pass for both pointer widths, including malformed requests, partial reads, repeated installation and callback exceptions. All four Memory32/Memory64 × 1/2-worker configurations pass in Node and Chrome after rebase. Fresh CI is queued.

@github-actions

github-actions Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

LLGo WebAssembly build benchmarks

249a8b39161c | workflow run | long-term charts

WebAssembly output sizes
Example, profile and compiler Wasm module vs base Generated JS glue vs base
cprintf/j32-emscripten/LLGo 154193 B 0 B / +0.0% 88742 B +13852 B / +18.5% (worse)
cprintf/j32-goos-js/LLGo 152295 B 0 B / +0.0% 73165 B 0 B / +0.0%
cprintf/j64-emscripten-memory64/LLGo 141248 B 0 B / +0.0% 92630 B +13852 B / +17.6% (worse)
cprintf/w32-goos-wasip1/LLGo 147689 B 0 B / +0.0% 0 B 0 B / 0.0%
cprintf/w32-wasi/LLGo 147797 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/j32-emscripten/LLGo 3213428 B 0 B / +0.0% 132442 B +13852 B / +11.7% (worse)
fmtprintf/j32-goos-js/Go 2526852 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/j32-goos-js/LLGo 3192523 B 0 B / +0.0% 101635 B 0 B / +0.0%
fmtprintf/j64-emscripten-memory64/LLGo 2950033 B 0 B / +0.0% 139285 B +13852 B / +11.0% (worse)
fmtprintf/w32-goos-wasip1/Go 2500019 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/w32-goos-wasip1/LLGo 2844343 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/w32-wasi/LLGo 2711131 B 0 B / +0.0% 0 B 0 B / 0.0%
j32-emscripten/LLGo 153428 B 0 B / +0.0% 88742 B +13852 B / +18.5% (worse)
j32-goos-js/Go 1895533 B 0 B / +0.0% 0 B 0 B / 0.0%
j32-goos-js/LLGo 151768 B 0 B / +0.0% 73165 B 0 B / +0.0%
j64-emscripten-memory64/LLGo 140582 B 0 B / +0.0% 92630 B +13852 B / +17.6% (worse)
reflectcall/j32-emscripten/LLGo 1543771 B 0 B / +0.0% 105908 B +13852 B / +15.0% (worse)
reflectcall/j32-goos-js/Go 2191221 B 0 B / +0.0% 0 B 0 B / 0.0%
reflectcall/j32-goos-js/LLGo 1546222 B 0 B / +0.0% 90331 B 0 B / +0.0%
reflectcall/j64-emscripten-memory64/LLGo 1428426 B 0 B / +0.0% 111641 B +13852 B / +14.2% (worse)
reflectcall/w32-goos-wasip1/Go 2205707 B 0 B / +0.0% 0 B 0 B / 0.0%
reflectcall/w32-goos-wasip1/LLGo 1548815 B 0 B / +0.0% 0 B 0 B / 0.0%
reflectcall/w32-wasi/LLGo 1474699 B 0 B / +0.0% 0 B 0 B / 0.0%
w32-goos-wasip1/Go 1909947 B 0 B / +0.0% 0 B 0 B / 0.0%
w32-goos-wasip1/LLGo 147023 B 0 B / +0.0% 0 B 0 B / 0.0%
w32-wasi/LLGo 147197 B 0 B / +0.0% 0 B 0 B / 0.0%
LLGo WebAssembly build measurements
Example and profile Build vs base
j32-emscripten 6.170 s +93.09 ms / +1.5% (worse)
j32-goos-js 5.927 s -54.29 ms / -0.9% (better)
j64-emscripten-memory64 5.221 s +125.6 ms / +2.5% (worse)
reflectcall/w32-wasi 25.996 s +154.1 ms / +0.6% (worse)
w32-goos-wasip1 4.661 s +84.56 ms / +1.8% (worse)
w32-wasi 4.505 s +36.37 ms / +0.8% (worse)

Compared with abe36f633097 measured in the same runner job.

@cpunion
cpunion force-pushed the codex/browser-fs-workers branch 3 times, most recently from 246d737 to c627831 Compare September 29, 2026 11:34

@visualfc visualfc left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Non-blocking follow-up: the runtime-thread HEAPU8 view can go stale after a worker-side payload malloc grows WebAssembly.Memory. Fine to merge as-is; current Node/Chrome fixtures use small buffers and would not hit this.

Comment thread internal/wasmworkers/browser_fs.js
@cpunion
cpunion force-pushed the codex/browser-fs-workers branch from c627831 to 249a8b3 Compare September 30, 2026 05:48
@cpunion
cpunion merged commit 4d51d8e into xgo-dev:main Sep 30, 2026
84 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants