Support GCC and Clang PGO builds with embedded profile flushing - #1
Draft
LauJosefsen wants to merge 1 commit into
Draft
LauJosefsen wants to merge 1 commit into
LauJosefsen wants to merge 1 commit into
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The existing PGO targets assume GCC-style profiles and can reuse objects compiled for a different build phase. Embedded hosts that bypass C runtime exit handlers can also finish training without writing their counters.
This change makes the optional PGO workflow support GCC and Clang, rebuilds objects for each phase, checks for training data before replacing binaries, and merges Clang profiles before profile use. Generation uses atomic counters and flushes profiles at PHP module shutdown. The shutdown function is excluded from profiling in both phases to avoid a profile-shape mismatch around the instrumentation-only flush hook.
Includes documentation for representative training and a synthetic Linux shared-embed smoke test covering build transitions, missing profiles, flushing before
_Exit(), and identical output without instrumentation after profile use. Normal build flags remain unchanged, and the workflow does not enable JIT.Benchmark results
Exploratory source-build measurements on PHP 8.5.11 ZTS, Laravel 13.31.0, an AMD Ryzen 9 7950X, GCC 14.2, and Clang 23.1.2. JIT remained disabled throughout. These measurements predate this PR and evaluate the PGO build strategies; they are not measurements of the exact PR commit.
Median CPU reduction relative to the GCC hybrid-interpreter baseline (higher means less CPU):
The fresh-application setup uses FrankenPHP classic mode with warm shared OPcache; it is not an NTS PHP-FPM measurement. HTTP CPU includes the server and embedded PHP process. CLI job measurements cover Laravel bus dispatch and application work, excluding broker I/O.
Job-work detail, with the minimum/maximum paired reduction over three rounds in parentheses:
Method and scope
Tail calls alone roughly matched the GCC baseline in the short HTTP screening run and improved job CPU by about 3.3%. GCC LTO and O3 showed no consistent additional benefit over PGO. The GCC LTO experiment excluded two VM/helper translation units; Clang full/ThinLTO failed to link and was not benchmarked. The benchmarked PHP 8.5.11 source predates the GCC LTO register-reservation fix in GH-23602, merged on September 16, 2026; these measurements do not evaluate that fix.
These are short workstation measurements on x86-64 ZTS. ARM64 and NTS PHP-FPM have not been measured. SQLite and a local HTTP mock do not model production I/O latency, and the workload mix does not establish a universal application or fleet-wide gain.
Earlier work and overlap
PGO itself and the tail-call interpreter are existing techniques. This draft proposes build integration and reliability improvements, and should be assessed alongside the following earlier work:
prof-useleaving instrumented objects in place without an explicitprof-clean. The reporter described a 7% application performance improvement after successful training. Automatic cleanup was discussed, and Dmitry Stogov objected because it complicates selectively applying PGO to individual files. This draft's full phase rebuilds change that selective workflow; compatibility remains a design question before an upstream proposal.The potential contribution is a compiler-aware Unix workflow covering GCC and Clang, embedded profile persistence, guarded build transitions, and a regression smoke test. Neither PGO nor shutdown flushing should be presented as a new idea. Earlier work also does not establish that one training profile improves every application.