perf(https): re-arm TCP_QUICKACK around the TLS handshake to avoid the Nagle/delayed-ACK stall - #307
Merged
Merged
Conversation
…e Nagle/delayed-ACK stall Every HTTPS connection from a client without TCP_NODELAY (stock ab, many OpenSSL-based clients) paid a fixed ~40-50 ms before its first request was processed: the client holds the HTTP request behind Nagle until the server ACKs the client's TLS Finished, and the server's delayed ACK (re-entered during the handshake segment exchange; the one-shot TCP_QUICKACK set at accept time does not survive it) waits out the full delayed-ACK window. Re-arm TCP_QUICKACK through the handshake (blocking accept loop and shepherd) and once the session is established, so the Finished is ACKed immediately and the request arrives right behind the handshake. Measured on a 12-core Linux box, ECDSA P-256 cert, loopback: | benchmark | before | after | |------------------------------|---------|---------| | request RTT, sequential | 43 ms | 0.2 ms | | handshake+request, 1 client | 22/s | 855/s | | 8 clients | 175/s | 5596/s | | 32 clients | 705/s | 7199/s | | ab -n 3000 -c 8 | 176/s | 1674/s | | ab -n 10000 -c 32 | 686/s | 1688/s | | ab -n 8000 -c 32 (RSA-4096) | 610/s | 983/s | Controls unchanged: plaintext ab 14.1k/s, keep-alive wrk HTTPS 189.6k/s (vs 189.4k/s before). test_https, test_https_park, test_https_full_gc and test_http2 all pass. Fixes the dominant term behind the TLS handshake plateau in #306.
Member
Author
|
Nice. |
This was referenced Oct 4, 2026
gg582
added a commit
that referenced
this pull request
Oct 4, 2026
CI run 37201724744 (PR #309, shared AMD EPYC 9V45 runner) failed only http_big_gbps: 7.8 measured vs the 8.0 backstop, while every other gate passed with the local-box numbers as reference. The absolute backstops were calibrated on a quiet 12-core desktop and are too tight for noisy shared runners. Recalibrate each backstop against the observed CI values with variance headroom: tls_rtt_p50_ms <=5ms (unchanged; observed 0.070) https_churn_rps >=1200 -> >=1000 (observed 2013) https_keepalive_rps >=130k -> >=110k (observed 135570) https_big_gbps >=3 -> >=2.5 (observed 4.03) http_churn_rps >=8000 -> >=6000 (observed 15793) http_keepalive_rps >=180k -> >=150k (observed 206650) http_big_gbps >=8 -> >=5.5 (observed 7.8) Verified: a row rebuilt from the failed run's observed values now passes 7/7, and the pre-#307 row (43.2ms RTT p50, 683/s churn) still fails gates 1 and 2. The https_keepalive_ratio diagnostic signal is unchanged.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Fixes the dominant term behind the TLS handshake plateau reported in #306.
Every HTTPS connection from a client without TCP_NODELAY (stock
ab, many OpenSSL-based clients, several load generators) paid a fixed ~40–50 ms before its first request was processed:Finished, then holds the HTTP request behind Nagle until theFinishedis ACKed.TCP_QUICKACKset at accept time does not survive protocol processing) waits out the full ~40 ms delayed-ACK window.The kernel re-enters delayed-ACK (pingpong) mode while the TLS handshake segments flow, so a single
TCP_QUICKACKat accept time (already done inasync_server.c) does not cover the client's final flight.Fix: re-arm
TCP_QUICKACKthroughout the handshake — once beforeSSL_accept, after everySSL_acceptiteration (blocking pool loop incwist_https_acceptand the Linux shepherd inhttps_hs_shepherd), and once the session is established inhttps_wrap_established— so the client'sFinishedis ACKed immediately and the request arrives right behind the handshake. Best-effort only; on non-Linux or on setsockopt failure behavior is unchanged.A/B measurements
Same machine (12-core Linux, loopback), same bench server (
cwist_app_use_https,{"msg":"hello"}route), ECDSA P-256 cert unless noted. Before =dev@ 60c7bf2, after = this branch. Interleaved runs, raw binaries in /tmp.tlsbench t=1)tlsbench t=8)tlsbench t=32)ab -n 3000 -c 8(new conn per request)ab -n 10000 -c 32ab -n 8000 -c 32with RSA-4096 certControls (unchanged, no regression):
ab -n 20000 -c 32: 14,184/s before vs 14,115/s afterwrk -t4 -c100 -d10sHTTPS: 189.6k req/s before vs 189.4k afterNote: the earlier ~640 handshakes/s figure in #306 was this same stall (measured through clients that don't set
TCP_NODELAY); it was not a shard-count or crypto limit. The remaining gap to plaintext (~14k/s) is now genuine TLS+fork/accept cost, not ACK timing.Test evidence
make test_https— passmake test_https_park— passmake test_https_full_gc— passmake test_http2— pass (shareshttps_wrap_established)make libcwist.a— clean buildRisk
Low: the change only tunes ACK timing via a best-effort setsockopt, Linux-only, guarded by
defined(__linux__) && defined(TCP_QUICKACK). No behavior change on other platforms or on failure.Refs #306