Skip to content

plan9asm: support Uint128 signatures in modernc libc assembly wrappers - #2694

Draft
cpunion wants to merge 2 commits into
xgo-dev:mainfrom
cpunion:codex/fix-issue2691-struct-asm-20260929
Draft

cpunion wants to merge 2 commits into
xgo-dev:mainfrom
cpunion:codex/fix-issue2691-struct-asm-20260929

Conversation

@cpunion

@cpunion cpunion commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator

Refs #2691. This is a dependency-integration experiment, not a complete SQLite or ABI0 fix and not ready to merge.

Scope

  • Pin the flat-word-struct implementation from Infer flat word structs in Plan 9 asm signatures plan9asm#43 and test the reported modernc.org/libc Uint128 wrapper shape.
  • Verify LLVM module validity and both i64 field extractions from each Uint128 parameter using LLVM values and instructions, without depending on SSA names or printed IR formatting.
  • Document the effective fork dependency and its unchanged Apache-2.0 license in THIRD_PARTY_NOTICES.md.

Merge blockers

  • Remove the personal-fork replace and use an upstream plan9asm dependency containing the required support. The latest official release, v0.6.1, still fails this regression with unsupported struct type struct{Lo uint64; Hi uint64} when tested with a separate modfile. The replacement affects all LLGo builds, not only tests, and must not land on main.
  • Infer flat word structs in Plan 9 asm signatures plan9asm#43 is closed. The broader implementation staged for translator: expand multi-architecture coverage from ecosystem discovery plan9asm#40 already contains CLI aggregate support and backend ABI0 stack-call handling, but the audited library API still rejects non-empty structs and does not attach outgoing frame information to referenced Go callees. Reuse that implementation and close the API gap tracked by ABI0 wrapper CALL lowering ignores outgoing SP arguments plan9asm#44; do not maintain a separate partial lowering path.
  • Run the original SQLite example on Linux/amd64 after those fixes. The real modernc.org/libc@v1.75.7 assembly file includes complex/structured arguments and callback trampolines. Signature and IR verification alone do not establish correct execution.

Validation

  • go test ./internal/plan9asm -count=1
  • go test ./internal/plan9asm -race -count=1
  • go vet ./internal/plan9asm
  • Negative control with official plan9asm v0.6.1 reproduces the unsupported-struct error; no test skips were added.

@codecov

codecov Bot commented Sep 29, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@cpunion cpunion changed the title Support modernc.org/libc Uint128 ABI0 wrapper (#2691) WIP: Diagnose modernc.org/libc ABI0 wrappers (#2691) Sep 29, 2026
@github-actions

github-actions Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

LLGo baseline benchmarks

71d85c0f1081 | workflow run | long-term charts

Program measurements

Platform Workload File size vs base Text size vs base Build vs base Run vs base
Linux cprintf 7160 B 0 B / +0.0% 387 B 0 B / +0.0% 524.919 ms -115.5 ms / -18.0% (better) 1.279 ms -93.42 us / -6.8% (better)
Linux cprintf-lto 6912 B 0 B / +0.0% 368 B 0 B / +0.0% 553.590 ms -64.27 ms / -10.4% (better) 1.450 ms +181.3 us / +14.3% (worse)
Linux fmtprintf 1662272 B 0 B / +0.0% 498359 B 0 B / +0.0% 4.001 s -307.7 ms / -7.1% (better) 3.089 ms -388 us / -11.2% (better)
Linux fmtprintf-lto 1500384 B 0 B / +0.0% 436113 B 0 B / +0.0% 11.562 s +235 ms / +2.1% (worse) 2.989 ms -14.61 us / -0.5% (better)
Linux println 68992 B 0 B / +0.0% 16783 B 0 B / +0.0% 566.022 ms -67.1 ms / -10.6% (better) 1.772 ms +95.81 us / +5.7% (worse)
Linux println-lto 59848 B 0 B / +0.0% 14199 B 0 B / +0.0% 834.943 ms -50.45 ms / -5.7% (better) 1.649 ms -238.2 us / -12.6% (better)
macOS cprintf 68064 B 0 B / +0.0% 4429 B 0 B / +0.0% 669.213 ms +27.88 ms / +4.3% (worse) 2.261 ms -249 us / -9.9% (better)
macOS cprintf-lto 84288 B 0 B / +0.0% 193 B 0 B / +0.0% 723.140 ms +58.92 ms / +8.9% (worse) 2.542 ms -219.9 us / -8.0% (better)
macOS fmtprintf 1504672 B 0 B / +0.0% 874572 B 0 B / +0.0% 2.818 s +124.6 ms / +4.6% (worse) 3.991 ms -695.2 us / -14.8% (better)
macOS fmtprintf-lto 1192704 B 0 B / +0.0% 848216 B 0 B / +0.0% 6.608 s +259.5 ms / +4.1% (worse) 3.540 ms -641.7 us / -15.3% (better)
macOS println 117136 B 0 B / +0.0% 37421 B 0 B / +0.0% 629.686 ms -3.614 ms / -0.6% (better) 3.025 ms +16.04 us / +0.5% (worse)
macOS println-lto 119472 B 0 B / +0.0% 34824 B 0 B / +0.0% 859.309 ms +104.5 ms / +13.8% (worse) 3.863 ms +899.7 us / +30.4% (worse)
Windows MinGW cprintf 19456 B 0 B / +0.0% 4550 B 0 B / +0.0% 1.320 s +28.1 ms / +2.2% (worse) 3.785 ms -241.2 us / -6.0% (better)
Windows MinGW cprintf-lto 17920 B 0 B / +0.0% 4486 B 0 B / +0.0% 1.336 s -11.88 ms / -0.9% (better) 3.611 ms -129.9 us / -3.5% (better)
Windows MinGW fmtprintf 1932800 B 0 B / +0.0% 598950 B 0 B / +0.0% 4.304 s +32.34 ms / +0.8% (worse) 9.059 ms +595 us / +7.0% (worse)
Windows MinGW fmtprintf-lto 1957376 B 0 B / +0.0% 547318 B 0 B / +0.0% 10.435 s +224.3 ms / +2.2% (worse) 8.682 ms +157.2 us / +1.8% (worse)
Windows MinGW println 75776 B 0 B / +0.0% 25142 B 0 B / +0.0% 1.336 s +3.218 ms / +0.2% (worse) 7.313 ms +599 us / +8.9% (worse)
Windows MinGW println-lto 69120 B 0 B / +0.0% 21990 B 0 B / +0.0% 1.562 s +16.77 ms / +1.1% (worse) 7.770 ms +974.1 us / +14.3% (worse)
Windows MinGW 386 cprintf 43520 B 0 B / +0.0% 5326 B 0 B / +0.0% 1.367 s +73.71 ms / +5.7% (worse) 5.532 ms -228.6 us / -4.0% (better)
Windows MinGW 386 cprintf-lto 20992 B 0 B / +0.0% 5094 B 0 B / +0.0% 1.337 s -4.061 ms / -0.3% (better) 5.074 ms -590.1 us / -10.4% (better)
Windows MinGW 386 fmtprintf 1896448 B 0 B / +0.0% 472414 B 0 B / +0.0% 4.655 s +225.6 ms / +5.1% (worse) 12.380 ms +1.831 ms / +17.4% (worse)
Windows MinGW 386 fmtprintf-lto 2180608 B 0 B / +0.0% 451258 B 0 B / +0.0% 10.782 s +823.3 ms / +8.3% (worse) 14.298 ms +3.471 ms / +32.1% (worse)
Windows MinGW 386 println 96256 B 0 B / +0.0% 21458 B 0 B / +0.0% 1.406 s +59.48 ms / +4.4% (worse) 9.782 ms +247.3 us / +2.6% (worse)
Windows MinGW 386 println-lto 74240 B 0 B / +0.0% 19314 B 0 B / +0.0% 1.671 s +135.5 ms / +8.8% (worse) 11.405 ms +1.998 ms / +21.2% (worse)
Windows MinGW ARM64 cprintf 18944 B 0 B / +0.0% 4408 B 0 B / +0.0% 1.545 s +2.483 ms / +0.2% (worse) 6.392 ms -105.6 us / -1.6% (better)
Windows MinGW ARM64 cprintf-lto 17920 B 0 B / +0.0% 4340 B 0 B / +0.0% 1.589 s +21.52 ms / +1.4% (worse) 6.353 ms -212.5 us / -3.2% (better)
Windows MinGW ARM64 fmtprintf 1819136 B 0 B / +0.0% 509608 B 0 B / +0.0% 4.430 s +63.11 ms / +1.4% (worse) 12.120 ms -525.1 us / -4.2% (better)
Windows MinGW ARM64 fmtprintf-lto 1879040 B 0 B / +0.0% 476152 B 0 B / +0.0% 9.681 s -96 ms / -1.0% (better) 12.808 ms +139.8 us / +1.1% (worse)
Windows MinGW ARM64 println 72192 B 0 B / +0.0% 23876 B 0 B / +0.0% 1.543 s -19.45 ms / -1.2% (better) 11.128 ms -71 us / -0.6% (better)
Windows MinGW ARM64 println-lto 68608 B 0 B / +0.0% 21224 B 0 B / +0.0% 1.755 s -4.38 ms / -0.2% (better) 11.230 ms +64.2 us / +0.6% (worse)
Windows MSVC cprintf 120320 B 0 B / +0.0% 65798 B 0 B / +0.0% 1.171 s -196.5 ms / -14.4% (better) 3.877 ms +536.5 us / +16.1% (worse)
Windows MSVC cprintf-lto 119808 B 0 B / +0.0% 65734 B 0 B / +0.0% 1.196 s +48.62 ms / +4.2% (worse) 4.268 ms +861.2 us / +25.3% (worse)
Windows MSVC fmtprintf 1643008 B 0 B / +0.0% 694502 B 0 B / +0.0% 4.207 s -671.4 ms / -13.8% (better) 11.246 ms +1.081 ms / +10.6% (worse)
Windows MSVC fmtprintf-lto 1634816 B 0 B / +0.0% 647062 B 0 B / +0.0% 9.751 s +79.71 ms / +0.8% (worse) 11.449 ms +1.431 ms / +14.3% (worse)
Windows MSVC println 194560 B 0 B / +0.0% 120822 B 0 B / +0.0% 1.176 s +36.93 ms / +3.2% (worse) 8.286 ms +1.136 ms / +15.9% (worse)
Windows MSVC println-lto 192512 B 0 B / +0.0% 118342 B 0 B / +0.0% 1.360 s +30.11 ms / +2.3% (worse) 7.613 ms -102.9 us / -1.3% (better)
Windows MSVC 386 cprintf 9728 B 0 B / +0.0% 3931 B 0 B / +0.0% 1.042 s -157.2 ms / -13.1% (better) 5.143 ms -3.348 ms / -39.4% (better)
Windows MSVC 386 cprintf-lto 9216 B 0 B / +0.0% 3853 B 0 B / +0.0% 1.083 s +24.03 ms / +2.3% (worse) 5.592 ms -4.6 us / -0.1% (better)
Windows MSVC 386 fmtprintf 1204224 B 0 B / +0.0% 455804 B 0 B / +0.0% 4.319 s +451.9 ms / +11.7% (worse) 12.473 ms +1.06 ms / +9.3% (worse)
Windows MSVC 386 fmtprintf-lto 1241088 B 0 B / +0.0% 427195 B 0 B / +0.0% 9.594 s +359.3 ms / +3.9% (worse) 14.488 ms +591.1 us / +4.3% (worse)
Windows MSVC 386 println 36352 B 0 B / +0.0% 20324 B 0 B / +0.0% 1.095 s +24.02 ms / +2.2% (worse) 9.751 ms +291.6 us / +3.1% (worse)
Windows MSVC 386 println-lto 35840 B 0 B / +0.0% 18549 B 0 B / +0.0% 1.272 s +22.01 ms / +1.8% (worse) 9.541 ms +315.1 us / +3.4% (worse)
Windows MSVC ARM64 cprintf 11776 B 0 B / +0.0% 4192 B 0 B / +0.0% 1.253 s +48.97 ms / +4.1% (worse) 6.926 ms +340.3 us / +5.2% (worse)
Windows MSVC ARM64 cprintf-lto 10752 B 0 B / +0.0% 4084 B 0 B / +0.0% 1.262 s +26.6 ms / +2.2% (worse) 6.875 ms +204 us / +3.1% (worse)
Windows MSVC ARM64 fmtprintf 1386496 B 0 B / +0.0% 509544 B 0 B / +0.0% 3.979 s +136.8 ms / +3.6% (worse) 14.244 ms +925.2 us / +6.9% (worse)
Windows MSVC ARM64 fmtprintf-lto 1404928 B 0 B / +0.0% 476820 B 0 B / +0.0% 8.873 s -94.52 ms / -1.1% (better) 12.962 ms -1.254 ms / -8.8% (better)
Windows MSVC ARM64 println 45056 B 0 B / +0.0% 23908 B 0 B / +0.0% 1.262 s +55.62 ms / +4.6% (worse) 12.389 ms +305.1 us / +2.5% (worse)
Windows MSVC ARM64 println-lto 42496 B 0 B / +0.0% 21380 B 0 B / +0.0% 1.436 s +39.96 ms / +2.9% (worse) 11.826 ms +168.7 us / +1.4% (worse)
Core language and compiler benchmarks
Platform Benchmark ns/op vs base
Linux BenchmarkLookupPCRandom 14.680 ns/op +0.05 ns/op / +0.3% (worse)
Linux BenchmarkMergeCompilerFlags 223.600 ns/op +29.3 ns/op / +15.1% (worse)
Linux BenchmarkMergeLinkerFlags 151 ns/op +24.3 ns/op / +19.2% (worse)
Linux BenchmarkChannelBuffered 55.520 ns/op -0.54 ns/op / -1.0% (better)
Linux BenchmarkChannelHandoff 15090 ns/op +1319 ns/op / +9.6% (worse)
Linux BenchmarkDefer 67.280 ns/op +9.11 ns/op / +15.7% (worse)
Linux BenchmarkDirectCall 1.609 ns/op +0.047 ns/op / +3.0% (worse)
Linux BenchmarkGlobalRead 1.167 ns/op -0.008 ns/op / -0.7% (better)
Linux BenchmarkGlobalWrite 8.099 ns/op +0.322 ns/op / +4.1% (worse)
Linux BenchmarkGoroutine 33446 ns/op +9432 ns/op / +39.3% (worse)
Linux BenchmarkInterfaceCall 6.001 ns/op +0.065 ns/op / +1.1% (worse)
Linux BenchmarkRuntimeGetG 3.028 ns/op -0.028 ns/op / -0.9% (better)
macOS BenchmarkLookupPCRandom 13.090 ns/op +0.32 ns/op / +2.5% (worse)
macOS BenchmarkMergeCompilerFlags 93.780 ns/op -13.72 ns/op / -12.8% (better)
macOS BenchmarkMergeLinkerFlags 68.930 ns/op -2.03 ns/op / -2.9% (better)
macOS BenchmarkChannelBuffered 26.490 ns/op +0.95 ns/op / +3.7% (worse)
macOS BenchmarkChannelHandoff 9377 ns/op +1212 ns/op / +14.8% (worse)
macOS BenchmarkDefer 32.380 ns/op -1.84 ns/op / -5.4% (better)
macOS BenchmarkDirectCall 1.172 ns/op -0.063 ns/op / -5.1% (better)
macOS BenchmarkGlobalRead 1.214 ns/op +0.074 ns/op / +6.5% (worse)
macOS BenchmarkGlobalWrite 1.062 ns/op -0.103 ns/op / -8.8% (better)
macOS BenchmarkGoroutine 57379 ns/op +16917 ns/op / +41.8% (worse)
macOS BenchmarkInterfaceCall 4.028 ns/op +0.024 ns/op / +0.6% (worse)
macOS BenchmarkRuntimeGetG 2.281 ns/op +0.158 ns/op / +7.4% (worse)
Windows MinGW BenchmarkLookupPCRandom 13.140 ns/op -0.05 ns/op / -0.4% (better)
Windows MinGW BenchmarkMergeCompilerFlags 660.300 ns/op +36.9 ns/op / +5.9% (worse)
Windows MinGW BenchmarkMergeLinkerFlags 591.400 ns/op +51.6 ns/op / +9.6% (worse)
Windows MinGW BenchmarkChannelBuffered 30.780 ns/op +0.92 ns/op / +3.1% (worse)
Windows MinGW BenchmarkChannelHandoff 919.500 ns/op +64.9 ns/op / +7.6% (worse)
Windows MinGW BenchmarkDefer 59.570 ns/op -0.01 ns/op / -0.01678% (better)
Windows MinGW BenchmarkDirectCall 1.550 ns/op +0.002 ns/op / +0.1% (worse)
Windows MinGW BenchmarkGlobalRead 1.862 ns/op +0.004 ns/op / +0.2% (worse)
Windows MinGW BenchmarkGlobalWrite 2.465 ns/op +0.005 ns/op / +0.2% (worse)
Windows MinGW BenchmarkGoroutine 95272 ns/op +1876 ns/op / +2.0% (worse)
Windows MinGW BenchmarkInterfaceCall 8.369 ns/op -0.024 ns/op / -0.3% (better)
Windows MinGW BenchmarkRuntimeGetG 2.170 ns/op -0.003 ns/op / -0.1% (better)
Windows MinGW 386 BenchmarkLookupPCRandom 26.770 ns/op -0.05 ns/op / -0.2% (better)
Windows MinGW 386 BenchmarkMergeCompilerFlags 780.100 ns/op +14.5 ns/op / +1.9% (worse)
Windows MinGW 386 BenchmarkMergeLinkerFlags 696.300 ns/op +13.5 ns/op / +2.0% (worse)
Windows MinGW 386 BenchmarkChannelBuffered 39.460 ns/op +0.16 ns/op / +0.4% (worse)
Windows MinGW 386 BenchmarkChannelHandoff 918.700 ns/op +72.9 ns/op / +8.6% (worse)
Windows MinGW 386 BenchmarkDefer 42.600 ns/op -0.33 ns/op / -0.8% (better)
Windows MinGW 386 BenchmarkDirectCall 1.548 ns/op -0.002 ns/op / -0.1% (better)
Windows MinGW 386 BenchmarkGlobalRead 1.550 ns/op -0.007 ns/op / -0.4% (better)
Windows MinGW 386 BenchmarkGlobalWrite 7.780 ns/op -0.008 ns/op / -0.1% (better)
Windows MinGW 386 BenchmarkGoroutine 108372 ns/op -91 ns/op / -0.1% (better)
Windows MinGW 386 BenchmarkInterfaceCall 8.391 ns/op -0.001 ns/op / -0.01192% (better)
Windows MinGW 386 BenchmarkRuntimeGetG 2.171 ns/op -0.005 ns/op / -0.2% (better)
Windows MinGW ARM64 BenchmarkLookupPCRandom 11.990 ns/op -0.06 ns/op / -0.5% (better)
Windows MinGW ARM64 BenchmarkMergeCompilerFlags 582.400 ns/op +10.5 ns/op / +1.8% (worse)
Windows MinGW ARM64 BenchmarkMergeLinkerFlags 545.200 ns/op +13.1 ns/op / +2.5% (worse)
Windows MinGW ARM64 BenchmarkChannelBuffered 38.190 ns/op -0.43 ns/op / -1.1% (better)
Windows MinGW ARM64 BenchmarkChannelHandoff 1831 ns/op -188 ns/op / -9.3% (better)
Windows MinGW ARM64 BenchmarkDefer 57.110 ns/op +1.46 ns/op / +2.6% (worse)
Windows MinGW ARM64 BenchmarkDirectCall 0.590 ns/op 0 ns/op / +0.0%
Windows MinGW ARM64 BenchmarkGlobalRead 0.663 ns/op 0 ns/op / +0.0%
Windows MinGW ARM64 BenchmarkGlobalWrite 0.664 ns/op +0.0001 ns/op / +0.01507% (worse)
Windows MinGW ARM64 BenchmarkGoroutine 62191 ns/op +1974 ns/op / +3.3% (worse)
Windows MinGW ARM64 BenchmarkInterfaceCall 4.140 ns/op -0.001 ns/op / -0.02415% (better)
Windows MinGW ARM64 BenchmarkRuntimeGetG 1.769 ns/op -0.035 ns/op / -1.9% (better)
Windows MSVC BenchmarkLookupPCRandom 13.060 ns/op -0.18 ns/op / -1.4% (better)
Windows MSVC BenchmarkMergeCompilerFlags 636.300 ns/op +23.2 ns/op / +3.8% (worse)
Windows MSVC BenchmarkMergeLinkerFlags 560 ns/op +12.8 ns/op / +2.3% (worse)
Windows MSVC BenchmarkChannelBuffered 31.390 ns/op -0.75 ns/op / -2.3% (better)
Windows MSVC BenchmarkChannelHandoff 1073 ns/op -22 ns/op / -2.0% (better)
Windows MSVC BenchmarkDefer 54.220 ns/op -0.54 ns/op / -1.0% (better)
Windows MSVC BenchmarkDirectCall 1.549 ns/op +0.001 ns/op / +0.1% (worse)
Windows MSVC BenchmarkGlobalRead 1.549 ns/op -0.004 ns/op / -0.3% (better)
Windows MSVC BenchmarkGlobalWrite 2.474 ns/op +0.001 ns/op / +0.04044% (worse)
Windows MSVC BenchmarkGoroutine 91054 ns/op +1978 ns/op / +2.2% (worse)
Windows MSVC BenchmarkInterfaceCall 8.369 ns/op +0.003 ns/op / +0.03586% (worse)
Windows MSVC BenchmarkRuntimeGetG 1.860 ns/op +0.002 ns/op / +0.1% (worse)
Windows MSVC 386 BenchmarkLookupPCRandom 27.850 ns/op 0 ns/op / +0.0%
Windows MSVC 386 BenchmarkMergeCompilerFlags 764.200 ns/op -0.6 ns/op / -0.1% (better)
Windows MSVC 386 BenchmarkMergeLinkerFlags 720 ns/op +40.8 ns/op / +6.0% (worse)
Windows MSVC 386 BenchmarkChannelBuffered 43.710 ns/op +0.04 ns/op / +0.1% (worse)
Windows MSVC 386 BenchmarkChannelHandoff 818.400 ns/op -14.7 ns/op / -1.8% (better)
Windows MSVC 386 BenchmarkDefer 52.500 ns/op +1.04 ns/op / +2.0% (worse)
Windows MSVC 386 BenchmarkDirectCall 1.747 ns/op 0 ns/op / +0.0%
Windows MSVC 386 BenchmarkGlobalRead 1.973 ns/op +0.223 ns/op / +12.7% (worse)
Windows MSVC 386 BenchmarkGlobalWrite 8.980 ns/op -0.009 ns/op / -0.1% (better)
Windows MSVC 386 BenchmarkGoroutine 95366 ns/op +1352 ns/op / +1.4% (worse)
Windows MSVC 386 BenchmarkInterfaceCall 9.442 ns/op -0.009 ns/op / -0.1% (better)
Windows MSVC 386 BenchmarkRuntimeGetG 2.451 ns/op +0.006 ns/op / +0.2% (worse)
Windows MSVC ARM64 BenchmarkLookupPCRandom 12.310 ns/op +0.31 ns/op / +2.6% (worse)
Windows MSVC ARM64 BenchmarkMergeCompilerFlags 571.200 ns/op -8.9 ns/op / -1.5% (better)
Windows MSVC ARM64 BenchmarkMergeLinkerFlags 529.400 ns/op -9.9 ns/op / -1.8% (better)
Windows MSVC ARM64 BenchmarkChannelBuffered 37.660 ns/op -0.01 ns/op / -0.02655% (better)
Windows MSVC ARM64 BenchmarkChannelHandoff 1933 ns/op +36 ns/op / +1.9% (worse)
Windows MSVC ARM64 BenchmarkDefer 59.760 ns/op -1.81 ns/op / -2.9% (better)
Windows MSVC ARM64 BenchmarkDirectCall 0.589 ns/op -0.0001 ns/op / -0.01696% (better)
Windows MSVC ARM64 BenchmarkGlobalRead 0.664 ns/op 0 ns/op / +0.0%
Windows MSVC ARM64 BenchmarkGlobalWrite 3.754 ns/op +0.002 ns/op / +0.1% (worse)
Windows MSVC ARM64 BenchmarkGoroutine 58194 ns/op -2341 ns/op / -3.9% (better)
Windows MSVC ARM64 BenchmarkInterfaceCall 4.145 ns/op -0.001 ns/op / -0.02412% (better)
Windows MSVC ARM64 BenchmarkRuntimeGetG 1.803 ns/op +0.001 ns/op / +0.1% (worse)
Timer runtime benchmarks
Platform Operation and runtime ns/op vs base
Linux AfterFuncZeroDelivery/Go 911.300 ns/op +0.9 ns/op / +0.1% (worse)
Linux AfterFuncZeroDelivery/LLGo 45619 ns/op +7736 ns/op / +20.4% (worse)
Linux CreateStop/Go 302.200 ns/op +10.5 ns/op / +3.6% (worse)
Linux CreateStop/LLGo 1758 ns/op -227 ns/op / -11.4% (better)
Linux RearmStopped/Go 115.800 ns/op +1 ns/op / +0.9% (worse)
Linux RearmStopped/LLGo 1439 ns/op +19 ns/op / +1.3% (worse)
Linux ResetActive/Go 69.050 ns/op +1.59 ns/op / +2.4% (worse)
Linux ResetActive/LLGo 735 ns/op -106.2 ns/op / -12.6% (better)
Linux ResetHeap1024/Go 67.260 ns/op 0 ns/op / +0.0%
Linux ResetHeap1024/LLGo 186 ns/op -0.6 ns/op / -0.3% (better)
macOS AfterFuncZeroDelivery/Go 437.800 ns/op +7.4 ns/op / +1.7% (worse)
macOS AfterFuncZeroDelivery/LLGo 68863 ns/op +2631 ns/op / +4.0% (worse)
macOS CreateStop/Go 137.400 ns/op -36.7 ns/op / -21.1% (better)
macOS CreateStop/LLGo 496 ns/op +69.4 ns/op / +16.3% (worse)
macOS RearmStopped/Go 55.610 ns/op -2.81 ns/op / -4.8% (better)
macOS RearmStopped/LLGo 203.400 ns/op -146.9 ns/op / -41.9% (better)
macOS ResetActive/Go 42.250 ns/op -7.52 ns/op / -15.1% (better)
macOS ResetActive/LLGo 116.100 ns/op -56.2 ns/op / -32.6% (better)
macOS ResetHeap1024/Go 42.250 ns/op -0.4 ns/op / -0.9% (better)
macOS ResetHeap1024/LLGo 92.590 ns/op -2.64 ns/op / -2.8% (better)
Windows MinGW AfterFuncZeroDelivery/Go 548.800 ns/op -7.2 ns/op / -1.3% (better)
Windows MinGW AfterFuncZeroDelivery/LLGo 187443 ns/op -66 ns/op / -0.0352% (better)
Windows MinGW CreateStop/Go 115 ns/op -0.7 ns/op / -0.6% (better)
Windows MinGW CreateStop/LLGo 431.200 ns/op -1.5 ns/op / -0.3% (better)
Windows MinGW RearmStopped/Go 31.580 ns/op +0.05 ns/op / +0.2% (worse)
Windows MinGW RearmStopped/LLGo 282 ns/op +6.7 ns/op / +2.4% (worse)
Windows MinGW ResetActive/Go 20.100 ns/op -0.15 ns/op / -0.7% (better)
Windows MinGW ResetActive/LLGo 152.900 ns/op -14.9 ns/op / -8.9% (better)
Windows MinGW ResetHeap1024/Go 20.440 ns/op +0.06 ns/op / +0.3% (worse)
Windows MinGW ResetHeap1024/LLGo 123.700 ns/op -0.4 ns/op / -0.3% (better)
Windows MinGW 386 AfterFuncZeroDelivery/Go 971.900 ns/op +3.5 ns/op / +0.4% (worse)
Windows MinGW 386 AfterFuncZeroDelivery/LLGo 205360 ns/op +1959 ns/op / +1.0% (worse)
Windows MinGW 386 CreateStop/Go 191.800 ns/op -1.9 ns/op / -1.0% (better)
Windows MinGW 386 CreateStop/LLGo 484.200 ns/op -29.2 ns/op / -5.7% (better)
Windows MinGW 386 RearmStopped/Go 63.810 ns/op +0.37 ns/op / +0.6% (worse)
Windows MinGW 386 RearmStopped/LLGo 346.700 ns/op -0.3 ns/op / -0.1% (better)
Windows MinGW 386 ResetActive/Go 39.340 ns/op +0.21 ns/op / +0.5% (worse)
Windows MinGW 386 ResetActive/LLGo 988.500 ns/op +18 ns/op / +1.9% (worse)
Windows MinGW 386 ResetHeap1024/Go 39.590 ns/op +0.16 ns/op / +0.4% (worse)
Windows MinGW 386 ResetHeap1024/LLGo 188.600 ns/op +2.1 ns/op / +1.1% (worse)
Windows MinGW ARM64 AfterFuncZeroDelivery/Go 666 ns/op +5.4 ns/op / +0.8% (worse)
Windows MinGW ARM64 AfterFuncZeroDelivery/LLGo 142818 ns/op -7948 ns/op / -5.3% (better)
Windows MinGW ARM64 CreateStop/Go 194.600 ns/op -3.3 ns/op / -1.7% (better)
Windows MinGW ARM64 CreateStop/LLGo 389.700 ns/op -5.8 ns/op / -1.5% (better)
Windows MinGW ARM64 RearmStopped/Go 70.550 ns/op 0 ns/op / +0.0%
Windows MinGW ARM64 RearmStopped/LLGo 257.600 ns/op -1.6 ns/op / -0.6% (better)
Windows MinGW ARM64 ResetActive/Go 31.100 ns/op -0.04 ns/op / -0.1% (better)
Windows MinGW ARM64 ResetActive/LLGo 145.800 ns/op +1.8 ns/op / +1.3% (worse)
Windows MinGW ARM64 ResetHeap1024/Go 31.090 ns/op -0.02 ns/op / -0.1% (better)
Windows MinGW ARM64 ResetHeap1024/LLGo 127 ns/op -0.7 ns/op / -0.5% (better)
Windows MSVC AfterFuncZeroDelivery/Go 572.200 ns/op +8 ns/op / +1.4% (worse)
Windows MSVC AfterFuncZeroDelivery/LLGo 178215 ns/op -706 ns/op / -0.4% (better)
Windows MSVC CreateStop/Go 129.200 ns/op +13.2 ns/op / +11.4% (worse)
Windows MSVC CreateStop/LLGo 424.700 ns/op +0.6 ns/op / +0.1% (worse)
Windows MSVC RearmStopped/Go 31.180 ns/op -0.34 ns/op / -1.1% (better)
Windows MSVC RearmStopped/LLGo 256.600 ns/op -3.4 ns/op / -1.3% (better)
Windows MSVC ResetActive/Go 20.100 ns/op +0.03 ns/op / +0.1% (worse)
Windows MSVC ResetActive/LLGo 149.200 ns/op +7.8 ns/op / +5.5% (worse)
Windows MSVC ResetHeap1024/Go 20.380 ns/op -0.05 ns/op / -0.2% (better)
Windows MSVC ResetHeap1024/LLGo 126.500 ns/op +1.8 ns/op / +1.4% (worse)
Windows MSVC 386 AfterFuncZeroDelivery/Go 1006 ns/op +7.1 ns/op / +0.7% (worse)
Windows MSVC 386 AfterFuncZeroDelivery/LLGo 172062 ns/op +4099 ns/op / +2.4% (worse)
Windows MSVC 386 CreateStop/Go 220.700 ns/op +6.1 ns/op / +2.8% (worse)
Windows MSVC 386 CreateStop/LLGo 486.900 ns/op +8.6 ns/op / +1.8% (worse)
Windows MSVC 386 RearmStopped/Go 72.980 ns/op -0.23 ns/op / -0.3% (better)
Windows MSVC 386 RearmStopped/LLGo 336.300 ns/op +1.7 ns/op / +0.5% (worse)
Windows MSVC 386 ResetActive/Go 42.080 ns/op +0.13 ns/op / +0.3% (worse)
Windows MSVC 386 ResetActive/LLGo 1050 ns/op -6 ns/op / -0.6% (better)
Windows MSVC 386 ResetHeap1024/Go 42.330 ns/op -0.03 ns/op / -0.1% (better)
Windows MSVC 386 ResetHeap1024/LLGo 182.900 ns/op +0.3 ns/op / +0.2% (worse)
Windows MSVC ARM64 AfterFuncZeroDelivery/Go 664 ns/op -4.4 ns/op / -0.7% (better)
Windows MSVC ARM64 AfterFuncZeroDelivery/LLGo 139086 ns/op +535 ns/op / +0.4% (worse)
Windows MSVC ARM64 CreateStop/Go 199.400 ns/op -5.3 ns/op / -2.6% (better)
Windows MSVC ARM64 CreateStop/LLGo 436.600 ns/op -5.5 ns/op / -1.2% (better)
Windows MSVC ARM64 RearmStopped/Go 70.600 ns/op -0.04 ns/op / -0.1% (better)
Windows MSVC ARM64 RearmStopped/LLGo 274.900 ns/op +1.1 ns/op / +0.4% (worse)
Windows MSVC ARM64 ResetActive/Go 31.110 ns/op +0.13 ns/op / +0.4% (worse)
Windows MSVC ARM64 ResetActive/LLGo 153.300 ns/op -3.6 ns/op / -2.3% (better)
Windows MSVC ARM64 ResetHeap1024/Go 31.090 ns/op -0.06 ns/op / -0.2% (better)
Windows MSVC ARM64 ResetHeap1024/LLGo 137.400 ns/op -2.2 ns/op / -1.6% (better)

Compared with ef3fce46c9b3 measured in the same runner job.

@cpunion
cpunion marked this pull request as ready for review September 30, 2026 07:50
@cpunion cpunion changed the title WIP: Diagnose modernc.org/libc ABI0 wrappers (#2691) plan9asm: support Uint128 signatures in modernc libc assembly wrappers Sep 30, 2026

@fennoai fennoai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review summary

PR #2694 ("WIP: Diagnose modernc.org/libc ABI0 wrappers", #2691) adds one translation test plus a replace that redirects github.com/xgo-dev/plan9asm to a personal fork. The test is well-structured and consistent with the package's existing tests (mustTestPackage, TranslateSourceModuleForPkg, defer tr.Module.Dispose(), llvm.VerifyModule), and the ABI0 frame arithmetic ($56-52) is internally consistent. The two concerns are the committed fork redirect and a brittle/under-specified IR assertion — both inline below.

Findings: 1 merge-relevant (go.mod fork replace), 2 test-quality. No performance issues found.

Not placeable inline (outside the diff): THIRD_PARTY_NOTICES.md:70 still lists only github.com/xgo-dev/plan9asm, while the effective module graph (go.sum) now builds github.com/cpunion/plan9asm — the provenance notice no longer matches what ships.

Not buildable locally (LLVM headers unavailable), so this was judged by reading, not by running.

Comment thread go.mod

replace github.com/xgo-dev/llgo/runtime => ./runtime

replace github.com/xgo-dev/plan9asm => github.com/cpunion/plan9asm v0.0.0-20260929013439-34a93b2081cf

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 — a committed replace redirects a first-party compiler dependency to a personal fork. This shadows the github.com/xgo-dev/plan9asm v0.6.1 requirement for every build of the module (not only this test), moving control of compiled codegen from the xgo-dev org to an untagged personal-fork pseudo-version (v0.0.0-...). The shipped code differs from the org release — the module zip hashes differ (fork h1:6gLAPwaW... at go.sum:1 vs org v0.6.1 h1:WcwHdcNK... at go.sum:29) — so this is an unreviewed divergence on a compiler path, not a no-op. The replace also silently overrides any future go get -u of the real module until removed. For a WIP diagnostic, prefer a non-committed override (e.g. go.work) or upstream the fork and bump the require to a tagged release before merge; the fork redirect should not land in committed go.mod.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agreed: this replacement must not land on main, and I am leaving this P1 unresolved. I verified the latest official release is still v0.6.1 and reran the regression using a separate modfile without the fork override: it fails with unsupported struct type struct{Lo uint64; Hi uint64}. Removing the override now would break the regression; skipping it would hide the gap. The PR body now explicitly gates merge on an upstream dependency containing the fix, together with the API/ABI0 integration and execution tests. 71d85c0 also addresses the separate provenance finding in THIRD_PARTY_NOTICES.md; the pinned fork LICENSE is byte-identical to the bundled Plan9Asm license. This documentation does not resolve the dependency blocker.

Comment thread internal/plan9asm/uint128_asm_test.go Outdated
Comment thread internal/plan9asm/uint128_asm_test.go
@github-actions

Copy link
Copy Markdown

LLGo WebAssembly build benchmarks

71d85c0f1081 | workflow run | long-term charts

WebAssembly output sizes
Example, profile and compiler Wasm module vs base Generated JS glue vs base
cprintf/j32-emscripten/LLGo 146973 B 0 B / +0.0% 74890 B 0 B / +0.0%
cprintf/j32-goos-js/LLGo 145407 B 0 B / +0.0% 73165 B 0 B / +0.0%
cprintf/j64-emscripten-memory64/LLGo 134516 B 0 B / +0.0% 78778 B 0 B / +0.0%
cprintf/w32-goos-wasip1/LLGo 141956 B 0 B / +0.0% 0 B 0 B / 0.0%
cprintf/w32-wasi/LLGo 141704 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/j32-emscripten/LLGo 3203357 B 0 B / +0.0% 118556 B 0 B / +0.0%
fmtprintf/j32-goos-js/Go 2526852 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/j32-goos-js/LLGo 3182776 B 0 B / +0.0% 101635 B 0 B / +0.0%
fmtprintf/j64-emscripten-memory64/LLGo 2940299 B 0 B / +0.0% 125399 B 0 B / +0.0%
fmtprintf/w32-goos-wasip1/Go 2500019 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/w32-goos-wasip1/LLGo 2835788 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/w32-wasi/LLGo 2702221 B 0 B / +0.0% 0 B 0 B / 0.0%
j32-emscripten/LLGo 146205 B 0 B / +0.0% 74890 B 0 B / +0.0%
j32-goos-js/Go 1895533 B 0 B / +0.0% 0 B 0 B / 0.0%
j32-goos-js/LLGo 144876 B 0 B / +0.0% 73165 B 0 B / +0.0%
j64-emscripten-memory64/LLGo 133847 B 0 B / +0.0% 78778 B 0 B / +0.0%
reflectcall/j32-emscripten/LLGo 1534085 B 0 B / +0.0% 92056 B 0 B / +0.0%
reflectcall/j32-goos-js/Go 2191221 B 0 B / +0.0% 0 B 0 B / 0.0%
reflectcall/j32-goos-js/LLGo 1536830 B 0 B / +0.0% 90331 B 0 B / +0.0%
reflectcall/j64-emscripten-memory64/LLGo 1419257 B 0 B / +0.0% 97789 B 0 B / +0.0%
reflectcall/w32-goos-wasip1/Go 2205707 B 0 B / +0.0% 0 B 0 B / 0.0%
reflectcall/w32-goos-wasip1/LLGo 1540668 B 0 B / +0.0% 0 B 0 B / 0.0%
reflectcall/w32-wasi/LLGo 1466211 B 0 B / +0.0% 0 B 0 B / 0.0%
w32-goos-wasip1/Go 1909947 B 0 B / +0.0% 0 B 0 B / 0.0%
w32-goos-wasip1/LLGo 141170 B 0 B / +0.0% 0 B 0 B / 0.0%
w32-wasi/LLGo 140989 B 0 B / +0.0% 0 B 0 B / 0.0%
LLGo WebAssembly build measurements
Example and profile Build vs base
j32-emscripten 5.838 s -145.2 ms / -2.4% (better)
j32-goos-js 5.958 s -1.159 s / -16.3% (better)
j64-emscripten-memory64 5.191 s +53.11 ms / +1.0% (worse)
reflectcall/w32-wasi 26.312 s +444.6 ms / +1.7% (worse)
w32-goos-wasip1 4.689 s -283.7 ms / -5.7% (better)
w32-wasi 4.835 s +175.2 ms / +3.8% (worse)

Compared with ef3fce46c9b3 measured in the same runner job.

@cpunion
cpunion marked this pull request as draft October 1, 2026 02:30
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant