Skip to content

perf(https): re-arm TCP_QUICKACK around the TLS handshake to avoid the Nagle/delayed-ACK stall - #307

Merged
gg582 merged 1 commit into
devfrom
perf/https-quickack-handshake
Oct 4, 2026
Merged

gg582 merged 1 commit into
devfrom
perf/https-quickack-handshake

Conversation

@gg582

@gg582 gg582 commented Oct 4, 2026

Copy link
Copy Markdown
Member

Summary

Fixes the dominant term behind the TLS handshake plateau reported in #306.

Every HTTPS connection from a client without TCP_NODELAY (stock ab, many OpenSSL-based clients, several load generators) paid a fixed ~40–50 ms before its first request was processed:

  1. Client sends TLS Finished, then holds the HTTP request behind Nagle until the Finished is ACKed.
  2. The server's delayed ACK (re-entered during the handshake segment exchange — the one-shot TCP_QUICKACK set at accept time does not survive protocol processing) waits out the full ~40 ms delayed-ACK window.
  3. The request only arrives one delayed-ACK window after the handshake completes.

The kernel re-enters delayed-ACK (pingpong) mode while the TLS handshake segments flow, so a single TCP_QUICKACK at accept time (already done in async_server.c) does not cover the client's final flight.

Fix: re-arm TCP_QUICKACK throughout the handshake — once before SSL_accept, after every SSL_accept iteration (blocking pool loop in cwist_https_accept and the Linux shepherd in https_hs_shepherd), and once the session is established in https_wrap_established — so the client's Finished is ACKed immediately and the request arrives right behind the handshake. Best-effort only; on non-Linux or on setsockopt failure behavior is unchanged.

A/B measurements

Same machine (12-core Linux, loopback), same bench server (cwist_app_use_https, {"msg":"hello"} route), ECDSA P-256 cert unless noted. Before = dev @ 60c7bf2, after = this branch. Interleaved runs, raw binaries in /tmp.

Benchmark Before After Δ
Request RTT after handshake, sequential client 43 ms 0.2 ms -215x
Full conn (handshake+request), 1 client (tlsbench t=1) 22 conn/s 855 conn/s 39x
8 clients (tlsbench t=8) 175 conn/s 5,596 conn/s 32x
32 clients (tlsbench t=32) 705 conn/s 7,199 conn/s 10x
ab -n 3000 -c 8 (new conn per request) 176/s 1,674/s 9.5x
ab -n 10000 -c 32 686/s 1,688/s 2.5x
ab -n 8000 -c 32 with RSA-4096 cert 610/s 983/s 1.6x

Controls (unchanged, no regression):

  • Plaintext ab -n 20000 -c 32: 14,184/s before vs 14,115/s after
  • Keep-alive wrk -t4 -c100 -d10s HTTPS: 189.6k req/s before vs 189.4k after
  • Handshake crypto itself is unaffected (p50 ~0.8 ms before and after)

Note: the earlier ~640 handshakes/s figure in #306 was this same stall (measured through clients that don't set TCP_NODELAY); it was not a shard-count or crypto limit. The remaining gap to plaintext (~14k/s) is now genuine TLS+fork/accept cost, not ACK timing.

Test evidence

  • make test_https — pass
  • make test_https_park — pass
  • make test_https_full_gc — pass
  • make test_http2 — pass (shares https_wrap_established)
  • make libcwist.a — clean build

Risk

Low: the change only tunes ACK timing via a best-effort setsockopt, Linux-only, guarded by defined(__linux__) && defined(TCP_QUICKACK). No behavior change on other platforms or on failure.

Refs #306

…e Nagle/delayed-ACK stall

Every HTTPS connection from a client without TCP_NODELAY (stock ab, many
OpenSSL-based clients) paid a fixed ~40-50 ms before its first request was
processed: the client holds the HTTP request behind Nagle until the server
ACKs the client's TLS Finished, and the server's delayed ACK (re-entered
during the handshake segment exchange; the one-shot TCP_QUICKACK set at
accept time does not survive it) waits out the full delayed-ACK window.

Re-arm TCP_QUICKACK through the handshake (blocking accept loop and
shepherd) and once the session is established, so the Finished is ACKed
immediately and the request arrives right behind the handshake.

Measured on a 12-core Linux box, ECDSA P-256 cert, loopback:

| benchmark                    | before  | after   |
|------------------------------|---------|---------|
| request RTT, sequential      | 43 ms   | 0.2 ms  |
| handshake+request, 1 client  | 22/s    | 855/s   |
| 8 clients                    | 175/s   | 5596/s  |
| 32 clients                   | 705/s   | 7199/s  |
| ab -n 3000 -c 8              | 176/s   | 1674/s  |
| ab -n 10000 -c 32            | 686/s   | 1688/s  |
| ab -n 8000 -c 32 (RSA-4096)  | 610/s   | 983/s   |

Controls unchanged: plaintext ab 14.1k/s, keep-alive wrk HTTPS 189.6k/s
(vs 189.4k/s before). test_https, test_https_park, test_https_full_gc and
test_http2 all pass.

Fixes the dominant term behind the TLS handshake plateau in #306.
@gg582

gg582 commented Oct 4, 2026

Copy link
Copy Markdown
Member Author

Nice.

@gg582
gg582 merged commit f69c61b into dev Oct 4, 2026
16 checks passed
gg582 added a commit that referenced this pull request Oct 4, 2026
CI run 37201724744 (PR #309, shared AMD EPYC 9V45 runner) failed only
http_big_gbps: 7.8 measured vs the 8.0 backstop, while every other gate
passed with the local-box numbers as reference. The absolute backstops
were calibrated on a quiet 12-core desktop and are too tight for noisy
shared runners. Recalibrate each backstop against the observed CI values
with variance headroom:

  tls_rtt_p50_ms       <=5ms    (unchanged; observed 0.070)
  https_churn_rps      >=1200   -> >=1000   (observed 2013)
  https_keepalive_rps  >=130k   -> >=110k   (observed 135570)
  https_big_gbps       >=3      -> >=2.5    (observed 4.03)
  http_churn_rps       >=8000   -> >=6000   (observed 15793)
  http_keepalive_rps   >=180k   -> >=150k   (observed 206650)
  http_big_gbps        >=8      -> >=5.5    (observed 7.8)

Verified: a row rebuilt from the failed run's observed values now passes
7/7, and the pre-#307 row (43.2ms RTT p50, 683/s churn) still fails gates
1 and 2. The https_keepalive_ratio diagnostic signal is unchanged.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant