Skip to content

windows: qualify ARM64 and 386 MSVC targets (R10; depends on plan9asm#23) - #208

Closed
cpunion wants to merge 51 commits into
mainfrom
codex/windows-r10-msvc-arch-20260829
Closed

cpunion wants to merge 51 commits into
mainfrom
codex/windows-r10-msvc-arch-20260829

Conversation

@cpunion

@cpunion cpunion commented Aug 28, 2026 •

Copy link
Copy Markdown
Owner

Depends on xgo-dev/plan9asm#23.

Part of xgo-dev#2325.

This draft qualifies native Windows ARM64 and WoW64 386 MSVC targets while keeping the LLGo compiler, LLVM, LLDB, pkg-config, and other host tools on x64. GOOS=windows GOARCH=... remains authoritative: LLGo selects the canonical target triple and the architecture-matched Visual Studio, vcpkg, compiler-rt, Python, runtime/FFI, language, demo, benchmark, and scheduled GOROOT coverage. Explicit compiler and linker commands remain architecture-checked.

The implementation keeps three concerns separate:

  • x64 host tools are built before target activation;
  • target SDK and library selection follows GOARCH for amd64, arm64, and 386;
  • pure IR generation does not require a physical target linker, while native links do.

Windows/386 execution exposed and fixed target-specific compatibility gaps:

  • Go/386 structs and multi-result tuples use four-byte alignment for int64 and float64; Go callers and callees now share the same physical LLVM layout.
  • Native MSVC C structs retain their eight-byte alignment. C ABI lowering recursively bridges divergent Go and native aggregate layouts for parameters, results, exports, and callbacks.
  • Wide signed and unsigned indexes are checked before narrowing to 32-bit GEP indexes and use Go's standard high/low-word panic helpers.
  • Float-to-integer conversion follows the official Go 386 runtime semantics. Wide slice/channel lengths and capacities are range-checked before narrowing, including mixed signed/unsigned operand widths.
  • Pointer printing widens 32-bit values for the shared native formatting boundary.
  • The syscall bridge is marked SafeSEH-compatible, and function-entry/PC-line records use an explicit Win32 layout without conflating it with Go's 12-byte 386 struct layout.
  • Win32 stack walking skips LLGo wrapper frames, preserves exact function-entry records, and keeps Win64 entry slack isolated from 386.
  • Windows/386 CPU profiling now captures interrupted PCs and frame chains through the Win32 CONTEXT layout while retaining the existing Go runtime/pprof API.

Go 1.27 additionally moved the Windows ARM64 processor-feature query behind internal/cpu.isProcessorFeaturePresent. LLGo now supplies the same runtime hook as the official implementation through a thin Win32 IsProcessorFeaturePresent call, without adding a C wrapper.

The initial branch replaced five groups of Go 386 assembly with LLGo source patches. Those patches are no longer present. With plan9asm #23 (pinned to its current reviewed head), LLGo compiles the official Go internal/bytealg, internal/chacha8rand, internal/runtime/atomic, internal/runtime/syscall/windows, and math assembly. Only register-ABI metadata for two local internal/bytealg helpers remains LLGo-specific.

CI changes preserve main's current Go 1.20–1.27 coverage and add ARM64/386 rows to the existing Windows matrices. The compiler is always built with the current toolchain before activating a target architecture. Benchmark jobs compare base and head on the same runner; scheduled GOROOT reporting includes Windows MSVC amd64, ARM64, 386, and MinGW amd64 without multiplying old-toolchain Windows lanes.

Local validation on the rebased revision:

  • macOS/ARM64: go test ./ssa ./internal/cabi, go test ./internal/plan9asm with the plan9asm workspace, and go test -timeout=20m ./cl (621.6 s)
  • macOS/ARM64: all internal/build tests; the workspace-discovery test was also rerun with GOWORK unset as required by that test
  • Windows/386 MSVC: complete test/go execution with Go 1.27 and plan9asm feat(coro): add runnable stackless scheduler and panic prototype #23 current head
  • Windows/386 MSVC: runtime/pprof CPU start/stop, duplicate-start, goroutine sampling, and fault-recovery tests
  • Windows/ARM64 MSVC: runtime line/statement metadata, concurrent first-use function lookup, deferred-panic source line, and logical runtime-tail tests
  • Windows/386 MSVC: standard-library hash/crc32, math/bits, math/rand/v2, and full runtime/pprof suites before the final rebase; the affected profiling cases were rerun after it
  • Windows/386 and ARM64 outputs were identified as IMAGE_FILE_MACHINE_I386 and IMAGE_FILE_MACHINE_ARM64 and executed in the Windows 11 ARM64 VM

The fresh GitHub Actions matrix, Codecov result, and architecture-specific benchmark comparison are required before this PR is marked ready for review.

@cpunion cpunion added the windows Windows platform support label Aug 28, 2026
@github-actions

github-actions Bot commented Aug 28, 2026 •

Copy link
Copy Markdown

LLGo baseline benchmarks

8f9cd1de92d9 | workflow run | long-term charts

Program measurements

Platform Workload File size vs base Text size vs base Build vs base Run vs base
Linux cprintf 19560 B +64 B / +0.3% (worse) 387 B 0 B / +0.0% 347.796 ms +24.77 ms / +7.7% (worse) 1.263 ms -13.94 us / -1.1% (better)
Linux cprintf-lto 19392 B +64 B / +0.3% (worse) 368 B 0 B / +0.0% 342.562 ms +9.716 ms / +2.9% (worse) 1.265 ms +16.19 us / +1.3% (worse)
Linux fmtprintf 1608368 B +368 B / +0.02289% (worse) 490042 B 0 B / +0.0% 2.468 s +23.04 ms / +0.9% (worse) 3.171 ms +101.6 us / +3.3% (worse)
Linux fmtprintf-lto 1486096 B +352 B / +0.02369% (worse) 450888 B 0 B / +0.0% 8.171 s +29.89 ms / +0.4% (worse) 2.944 ms -74.32 us / -2.5% (better)
Linux println 62408 B +64 B / +0.1% (worse) 15286 B 0 B / +0.0% 335.375 ms +4.542 ms / +1.4% (worse) 1.573 ms +33.1 us / +2.1% (worse)
Linux println-lto 54224 B +64 B / +0.1% (worse) 12914 B 0 B / +0.0% 507.411 ms +3.759 ms / +0.7% (worse) 1.574 ms +13 us / +0.8% (worse)
macOS cprintf 84480 B 0 B / +0.0% 16797 B +64 B / +0.4% (worse) 764.998 ms +248.5 ms / +48.1% (worse) 6.979 ms +3.278 ms / +88.5% (worse)
macOS cprintf-lto 100704 B 0 B / +0.0% 16777 B +64 B / +0.4% (worse) 780.817 ms +189.5 ms / +32.1% (worse) 4.754 ms +242.3 us / +5.4% (worse)
macOS fmtprintf 1470304 B 0 B / +0.0% 867568 B +368 B / +0.04244% (worse) 2.932 s +205.6 ms / +7.5% (worse) 7.951 ms +1.317 ms / +19.8% (worse)
macOS fmtprintf-lto 1175536 B 0 B / +0.0% 863272 B +356 B / +0.04126% (worse) 8.947 s +1.134 s / +14.5% (worse) 8.561 ms -750 us / -8.1% (better)
macOS println 114784 B 0 B / +0.0% 35245 B +64 B / +0.2% (worse) 759.008 ms +288 ms / +61.1% (worse) 6.021 ms +716.5 us / +13.5% (worse)
macOS println-lto 118656 B 0 B / +0.0% 32897 B +64 B / +0.2% (worse) 949.349 ms +202.5 ms / +27.1% (worse) 7.648 ms +3.369 ms / +78.7% (worse)
Windows MinGW cprintf 20480 B 0 B / +0.0% 4662 B 0 B / +0.0% 845.851 ms +31.65 ms / +3.9% (worse) 4.654 ms +1.178 ms / +33.9% (worse)
Windows MinGW cprintf-lto 18432 B 0 B / +0.0% 4582 B 0 B / +0.0% 858.733 ms -19.16 ms / -2.2% (better) 3.536 ms +62.8 us / +1.8% (worse)
Windows MinGW fmtprintf 1937408 B +512 B / +0.02643% (worse) 597398 B 0 B / +0.0% 3.530 s -47.38 ms / -1.3% (better) 8.560 ms -241.6 us / -2.7% (better)
Windows MinGW fmtprintf-lto 1991168 B +512 B / +0.02572% (worse) 590870 B 0 B / +0.0% 9.352 s -47.22 ms / -0.5% (better) 8.380 ms -195 us / -2.3% (better)
Windows MinGW println 74240 B 0 B / +0.0% 25062 B 0 B / +0.0% 847.842 ms +23.08 ms / +2.8% (worse) 7.052 ms -348.1 us / -4.7% (better)
Windows MinGW println-lto 67584 B 0 B / +0.0% 21734 B 0 B / +0.0% 1.057 s +24.89 ms / +2.4% (worse) 7.005 ms +748.4 us / +12.0% (worse)
Windows MSVC cprintf 12288 B 0 B / +0.0% 4438 B 0 B / +0.0% 771.860 ms -153.9 ms / -16.6% (better) 3.514 ms -3.225 ms / -47.9% (better)
Windows MSVC cprintf-lto 11776 B 0 B / +0.0% 4278 B 0 B / +0.0% 939.393 ms +158.1 ms / +20.2% (worse) 3.543 ms +40.4 us / +1.2% (worse)
Windows MSVC fmtprintf 1476096 B +512 B / +0.0347% (worse) 596982 B 0 B / +0.0% 3.563 s +37.84 ms / +1.1% (worse) 10.105 ms +1.312 ms / +14.9% (worse)
Windows MSVC fmtprintf-lto 1523712 B +512 B / +0.03361% (worse) 597782 B 0 B / +0.0% 10.832 s -53.87 ms / -0.5% (better) 10.607 ms +532.7 us / +5.3% (worse)
Windows MSVC println 47104 B 0 B / +0.0% 25062 B 0 B / +0.0% 769.249 ms -5.598 ms / -0.7% (better) 6.976 ms -82.1 us / -1.2% (better)
Windows MSVC println-lto 44032 B 0 B / +0.0% 22102 B 0 B / +0.0% 1.030 s -2.109 ms / -0.2% (better) 6.976 ms +254.8 us / +3.8% (worse)
Windows MSVC 386 cprintf 9728 B new 3930 B new 868.747 ms new 6.136 ms new
Windows MSVC 386 cprintf-lto 9216 B new 3840 B new 999.460 ms new 7.744 ms new
Windows MSVC 386 fmtprintf 1195520 B new 456560 B new 3.736 s new 11.867 ms new
Windows MSVC 386 fmtprintf-lto 1259520 B new 454021 B new 10.311 s new 11.538 ms new
Windows MSVC 386 println 35328 B new 19744 B new 844.947 ms new 9.005 ms new
Windows MSVC 386 println-lto 33280 B new 17873 B new 1.094 s new 9.739 ms new
Windows MSVC ARM64 cprintf 11264 B new 3976 B new 1.825 s new 8.479 ms new
Windows MSVC ARM64 cprintf-lto 10752 B new 3844 B new 1.846 s new 8.079 ms new
Windows MSVC ARM64 fmtprintf 1370112 B new 510252 B new 6.083 s new 17.290 ms new
Windows MSVC ARM64 fmtprintf-lto 1419776 B new 508364 B new 18.853 s new 18.219 ms new
Windows MSVC ARM64 println 43008 B new 22856 B new 1.777 s new 14.086 ms new
Windows MSVC ARM64 println-lto 40960 B new 20656 B new 2.253 s new 14.031 ms new
Core language and compiler benchmarks
Platform Benchmark ns/op vs base
Linux BenchmarkLookupPCRandom 13.380 ns/op +0.06 ns/op / +0.5% (worse)
Linux BenchmarkMergeCompilerFlags 140.200 ns/op -0.1 ns/op / -0.1% (better)
Linux BenchmarkMergeLinkerFlags 88.750 ns/op -0.08 ns/op / -0.1% (better)
Linux BenchmarkChannelBuffered 35.200 ns/op -0.34 ns/op / -1.0% (better)
Linux BenchmarkChannelHandoff 27410 ns/op +857 ns/op / +3.2% (worse)
Linux BenchmarkDefer 46.040 ns/op -0.54 ns/op / -1.2% (better)
Linux BenchmarkDirectCall 1.556 ns/op 0 ns/op / +0.0%
Linux BenchmarkGlobalRead 1.869 ns/op +0.312 ns/op / +20.0% (worse)
Linux BenchmarkGlobalWrite 2.488 ns/op +0.003 ns/op / +0.1% (worse)
Linux BenchmarkGoroutine 31598 ns/op -277 ns/op / -0.9% (better)
Linux BenchmarkInterfaceCall 8.719 ns/op +0.63 ns/op / +7.8% (worse)
Linux BenchmarkRuntimeGetG 2.491 ns/op +0.001 ns/op / +0.04016% (worse)
macOS BenchmarkLookupPCRandom 15.090 ns/op +2.15 ns/op / +16.6% (worse)
macOS BenchmarkMergeCompilerFlags 125.400 ns/op -12.7 ns/op / -9.2% (better)
macOS BenchmarkMergeLinkerFlags 78.790 ns/op +4.7 ns/op / +6.3% (worse)
macOS BenchmarkChannelBuffered 26.300 ns/op -9.64 ns/op / -26.8% (better)
macOS BenchmarkChannelHandoff 8003 ns/op -2927 ns/op / -26.8% (better)
macOS BenchmarkDefer 36.970 ns/op -11.14 ns/op / -23.2% (better)
macOS BenchmarkDirectCall 1.158 ns/op -0.252 ns/op / -17.9% (better)
macOS BenchmarkGlobalRead 1.085 ns/op -0.221 ns/op / -16.9% (better)
macOS BenchmarkGlobalWrite 1.060 ns/op -0.388 ns/op / -26.8% (better)
macOS BenchmarkGoroutine 45343 ns/op +11613 ns/op / +34.4% (worse)
macOS BenchmarkInterfaceCall 4.438 ns/op -1.089 ns/op / -19.7% (better)
macOS BenchmarkRuntimeGetG 2.369 ns/op -0.39 ns/op / -14.1% (better)
Windows MinGW BenchmarkLookupPCRandom 13.140 ns/op +0.09 ns/op / +0.7% (worse)
Windows MinGW BenchmarkMergeCompilerFlags 665.100 ns/op +49.3 ns/op / +8.0% (worse)
Windows MinGW BenchmarkMergeLinkerFlags 579.700 ns/op +21 ns/op / +3.8% (worse)
Windows MinGW BenchmarkChannelBuffered 34.650 ns/op +0.17 ns/op / +0.5% (worse)
Windows MinGW BenchmarkChannelHandoff 910.200 ns/op +47.6 ns/op / +5.5% (worse)
Windows MinGW BenchmarkDefer 58.540 ns/op -0.08 ns/op / -0.1% (better)
Windows MinGW BenchmarkDirectCall 1.554 ns/op +0.008 ns/op / +0.5% (worse)
Windows MinGW BenchmarkGlobalRead 1.549 ns/op 0 ns/op / +0.0%
Windows MinGW BenchmarkGlobalWrite 2.471 ns/op +0.002 ns/op / +0.1% (worse)
Windows MinGW BenchmarkGoroutine 91380 ns/op +4158 ns/op / +4.8% (worse)
Windows MinGW BenchmarkInterfaceCall 9.302 ns/op +0.002 ns/op / +0.02151% (worse)
Windows MinGW BenchmarkRuntimeGetG 2.501 ns/op +0.024 ns/op / +1.0% (worse)
Windows MSVC BenchmarkLookupPCRandom 13.140 ns/op -0.11 ns/op / -0.8% (better)
Windows MSVC BenchmarkMergeCompilerFlags 611.800 ns/op -4.6 ns/op / -0.7% (better)
Windows MSVC BenchmarkMergeLinkerFlags 532.800 ns/op -3.5 ns/op / -0.7% (better)
Windows MSVC BenchmarkChannelBuffered 34.910 ns/op +0.02 ns/op / +0.1% (worse)
Windows MSVC BenchmarkChannelHandoff 1030 ns/op -64 ns/op / -5.9% (better)
Windows MSVC BenchmarkDefer 55.750 ns/op -0.64 ns/op / -1.1% (better)
Windows MSVC BenchmarkDirectCall 1.862 ns/op +0.001 ns/op / +0.1% (worse)
Windows MSVC BenchmarkGlobalRead 1.857 ns/op -0.004 ns/op / -0.2% (better)
Windows MSVC BenchmarkGlobalWrite 2.452 ns/op +0.001 ns/op / +0.0408% (worse)
Windows MSVC BenchmarkGoroutine 86460 ns/op -2401 ns/op / -2.7% (better)
Windows MSVC BenchmarkInterfaceCall 8.986 ns/op -0.011 ns/op / -0.1% (better)
Windows MSVC BenchmarkRuntimeGetG 2.171 ns/op +0.002 ns/op / +0.1% (worse)
Windows MSVC 386 BenchmarkLookupPCRandom 26.730 ns/op new
Windows MSVC 386 BenchmarkMergeCompilerFlags 765.100 ns/op new
Windows MSVC 386 BenchmarkMergeLinkerFlags 696.500 ns/op new
Windows MSVC 386 BenchmarkChannelBuffered 43.420 ns/op new
Windows MSVC 386 BenchmarkChannelHandoff 925.600 ns/op new
Windows MSVC 386 BenchmarkDefer 49.400 ns/op new
Windows MSVC 386 BenchmarkDirectCall 1.548 ns/op new
Windows MSVC 386 BenchmarkGlobalRead 1.548 ns/op new
Windows MSVC 386 BenchmarkGlobalWrite 7.771 ns/op new
Windows MSVC 386 BenchmarkGoroutine 93396 ns/op new
Windows MSVC 386 BenchmarkInterfaceCall 9.946 ns/op new
Windows MSVC 386 BenchmarkRuntimeGetG 2.168 ns/op new
Windows MSVC ARM64 BenchmarkLookupPCRandom 12.080 ns/op new
Windows MSVC ARM64 BenchmarkMergeCompilerFlags 565.200 ns/op new
Windows MSVC ARM64 BenchmarkMergeLinkerFlags 540.300 ns/op new
Windows MSVC ARM64 BenchmarkChannelBuffered 43.810 ns/op new
Windows MSVC ARM64 BenchmarkChannelHandoff 3033 ns/op new
Windows MSVC ARM64 BenchmarkDefer 70.280 ns/op new
Windows MSVC ARM64 BenchmarkDirectCall 0.590 ns/op new
Windows MSVC ARM64 BenchmarkGlobalRead 0.664 ns/op new
Windows MSVC ARM64 BenchmarkGlobalWrite 3.756 ns/op new
Windows MSVC ARM64 BenchmarkGoroutine 76819 ns/op new
Windows MSVC ARM64 BenchmarkInterfaceCall 4.718 ns/op new
Windows MSVC ARM64 BenchmarkRuntimeGetG 1.801 ns/op new

Compared with e637682c8772 measured in the same runner job. Platforms without a paired baseline are marked new.

@cpunion cpunion changed the title windows: qualify ARM64 and 386 MSVC targets (R10, depends on xgo-dev/llgo#2440) windows: qualify ARM64 and 386 MSVC targets (R10; depends on xgo-dev/llgo#2440 and plan9asm#23) Aug 29, 2026
@cpunion
cpunion force-pushed the codex/windows-r9-native-toolchain-20260827 branch from d813a25 to 0b66e0c Compare August 29, 2026 12:42
@cpunion
cpunion force-pushed the codex/windows-r10-msvc-arch-20260829 branch from 91ac3c4 to 83296dc Compare August 29, 2026 14:21
@cpunion cpunion changed the title windows: qualify ARM64 and 386 MSVC targets (R10; depends on xgo-dev/llgo#2440 and plan9asm#23) windows: qualify ARM64 and 386 MSVC targets (R10; depends on plan9asm#23) Aug 29, 2026
@cpunion
cpunion changed the base branch from codex/windows-r9-native-toolchain-20260827 to main August 29, 2026 14:21
@cpunion
cpunion changed the base branch from main to codex/windows-r9-native-toolchain-20260827 August 29, 2026 14:22
@cpunion
cpunion force-pushed the codex/windows-r10-msvc-arch-20260829 branch from c0db0ee to 25e15b6 Compare August 30, 2026 00:02
@cpunion
cpunion changed the base branch from codex/windows-r9-native-toolchain-20260827 to main August 30, 2026 00:02
@cpunion
cpunion changed the base branch from main to codex/windows-r9-main-base-20260829 August 30, 2026 00:05
@cpunion
cpunion force-pushed the codex/windows-r10-msvc-arch-20260829 branch from 3a0312b to 25e15b6 Compare August 30, 2026 02:50
@cpunion
cpunion changed the base branch from codex/windows-r9-main-base-20260829 to main August 30, 2026 02:56
@cpunion
cpunion force-pushed the codex/windows-r10-msvc-arch-20260829 branch from 25e15b6 to 8f9cd1d Compare August 30, 2026 07:10
@cpunion
cpunion force-pushed the codex/windows-r10-msvc-arch-20260829 branch from 8f9cd1d to 416b083 Compare August 30, 2026 13:53
@cpunion

cpunion commented Aug 30, 2026

Copy link
Copy Markdown
Owner Author

Replaced by xgo-dev#2454 after rebasing R10 onto current main and integrating the released plan9asm v0.5.1 x87 configuration.

@cpunion cpunion closed this Aug 30, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

windows Windows platform support

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant