Skip to content

Support GCC and Clang PGO builds with embedded profile flushing - #1

Draft
LauJosefsen wants to merge 1 commit into
PHP-8.5from
lejo/reproducible-pgo-builds
Draft

LauJosefsen wants to merge 1 commit into
PHP-8.5from
lejo/reproducible-pgo-builds

Conversation

@LauJosefsen

@LauJosefsen LauJosefsen commented Oct 2, 2026 •

Copy link
Copy Markdown
Owner

The existing PGO targets assume GCC-style profiles and can reuse objects compiled for a different build phase. Embedded hosts that bypass C runtime exit handlers can also finish training without writing their counters.

This change makes the optional PGO workflow support GCC and Clang, rebuilds objects for each phase, checks for training data before replacing binaries, and merges Clang profiles before profile use. Generation uses atomic counters and flushes profiles at PHP module shutdown. The shutdown function is excluded from profiling in both phases to avoid a profile-shape mismatch around the instrumentation-only flush hook.

Includes documentation for representative training and a synthetic Linux shared-embed smoke test covering build transitions, missing profiles, flushing before _Exit(), and identical output without instrumentation after profile use. Normal build flags remain unchanged, and the workflow does not enable JIT.

Benchmark results

Exploratory source-build measurements on PHP 8.5.11 ZTS, Laravel 13.31.0, an AMD Ryzen 9 7950X, GCC 14.2, and Clang 23.1.2. JIT remained disabled throughout. These measurements predate this PR and evaluate the PGO build strategies; they are not measurements of the exact PR commit.

Median CPU reduction relative to the GCC hybrid-interpreter baseline (higher means less CPU):

Build strategy Persistent HTTP workers Fresh Laravel application per request Persistent CLI job work
GCC + PGO 5.7% 5.7% 8.6%
Clang tail-call interpreter + PGO 6.7% 7.0% 10.3%

The fresh-application setup uses FrankenPHP classic mode with warm shared OPcache; it is not an NTS PHP-FPM measurement. HTTP CPU includes the server and embedded PHP process. CLI job measurements cover Laravel bus dispatch and application work, excluding broker I/O.

Job-work detail, with the minimum/maximum paired reduction over three rounds in parentheses:

Job type GCC + PGO Clang tail calls + PGO
Mixed and wildcard validation 8.4% (6.7–9.8%) 11.5% (10.5–12.3%)
Data transformation and collections 7.2% (6.5–7.3%) 9.6% (8.8–10.0%)
Model serialization and casts 8.7% (7.1–10.4%) 10.7% (10.3–10.9%)
Container resolution and events 8.9% (7.2–12.4%) 9.8% (9.6–10.3%)

Method and scope

  • Training covered 13 mixed Laravel workload families in both HTTP worker modes, plus 2,400 CLI operations across ten process startups. It included routing, validation, collections, data transformation, model casts, Blade, container/events, SQLite, local HTTP calls, and instrumentation. Evaluation used different inputs and additional request families.
  • HTTP evaluation covered 14 workloads, three runtime variants, both worker modes, and three randomized rounds: 252 main samples plus 18 Blade repeats, totaling about 2.46 million timed requests. Each sample ran for four seconds after warm-up. The summary uses the five-second-warm-up Blade repeat, which had zero timed cache misses, and the other 13 main workloads. Original Blade samples remain part of the recorded experiment.
  • Persistent job evaluation covered four job types, ten runtime variants, three randomized rounds, and 3,000 iterations per type: 360,000 timed job executions.
  • Summary percentages give each workload equal weight: first take each workload's median paired CPU reduction over the three rounds, then the median across workloads. They are not estimates weighted to a production traffic mix.
  • HTTP response checks and job checksums matched across runtimes. No HTTP or socket errors were recorded, and trace span counts remained consistent between runtime variants.

Tail calls alone roughly matched the GCC baseline in the short HTTP screening run and improved job CPU by about 3.3%. GCC LTO and O3 showed no consistent additional benefit over PGO. The GCC LTO experiment excluded two VM/helper translation units; Clang full/ThinLTO failed to link and was not benchmarked. The benchmarked PHP 8.5.11 source predates the GCC LTO register-reservation fix in GH-23602, merged on September 16, 2026; these measurements do not evaluate that fix.

These are short workstation measurements on x86-64 ZTS. ARM64 and NTS PHP-FPM have not been measured. SQLite and a local HTTP mock do not model production I/O latency, and the workload mix does not establish a universal application or fleet-wide gain.

Earlier work and overlap

PGO itself and the tail-call interpreter are existing techniques. This draft proposes build integration and reliability improvements, and should be assessed alongside the following earlier work:

  • PHP bug #73111, reported in 2016, describes prof-use leaving instrumented objects in place without an explicit prof-clean. The reporter described a 7% application performance improvement after successful training. Automatic cleanup was discussed, and Dmitry Stogov objected because it complicates selectively applying PGO to individual files. This draft's full phase rebuilds change that selective workflow; compatibility remains a design question before an upstream proposal.
  • GH-8284 was incorporated through commit c4a9d1e in April 2022. It removes GCC profiles produced by build-time PHP execution so they do not contaminate application training. This draft retains that behavior while extending it to Clang.
  • static-php-cli PR #1126, opened in April 2026, proposes Clang/Zig PGO, atomic counters, profile merging, and explicit PHP/FrankenPHP shutdown flushing because Go bypasses C runtime exit handlers. That proposal directly overlaps the embedded-host problem addressed here. Related work remains open in #1138 and #1228.
  • FrankenPHP issue #463 records an August 2025 GCC/libphp PGO experiment reporting about 15% higher Symfony worker throughput but roughly 7% worse CLI performance. This reinforces the need for mixed training and held-out evaluation. Separately, FrankenPHP PR #2361, merged in May 2026, adds Go-side PGO tooling; its Hello World result is a different workload and optimization target from the PHP CPU measurements above.
  • PHP PR #23992, opened September 29, 2026 and currently open, adds Windows Clang PGO support. This draft focuses on the Unix make targets and embedded PHP lifecycle.
  • The tail-call interpreter was incorporated in commit 73b98a38 in August 2025. Its upstream non-JIT Symfony measurements brought Clang to approximately GCC performance, consistent with tail calls alone providing little additional HTTP benefit against a GCC baseline in this experiment. PHP PR #10564 proposes LTO support and remains open; its discussion includes a configuration with a substantial performance regression. Together with the newer GCC LTO fix, this argues for measuring each toolchain and source revision rather than assuming LTO improves PHP.

The potential contribution is a compiler-aware Unix workflow covering GCC and Clang, embedded profile persistence, guarded build transitions, and a regression smoke test. Neither PGO nor shutdown flushing should be presented as a new idea. Earlier work also does not establish that one training profile improves every application.

@LauJosefsen LauJosefsen self-assigned this Oct 2, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant