Skip to content

fix: reliable CI measurements + docs rewritten with validated numbers - #17

Merged
github-actions[bot] merged 7 commits into
mainfrom
fix/ci-measurements-validated-docs
Sep 27, 2026
Merged

github-actions[bot] merged 7 commits into
mainfrom
fix/ci-measurements-validated-docs

Conversation

@KissPeter

@KissPeter KissPeter commented Sep 27, 2026 •

Copy link
Copy Markdown
Owner

Summary

Replaces measurement plumbing that was quietly producing corrupted numbers (transient nginx 502s, ab timeouts, pool-exhaustion 500s were mixed into averages) and rewrites the 4 affected doc pages with data from a single fully-green CI run (36235373898).

Root cause

The old harness trusted the 3-run average even when runs contained non-2xx responses or timed out. Under -c100 concurrent load:

  • nginx to gunicorn: backlog=128 + short timeouts produced intermittent 502s from nginx during 1MB load; run_test.sh also did not propagate the container's exit code, so a failed run still reported "success".
  • connection pool (pool=2): queued requests waited >60s (httpx pool timeout) and returned 500s on /sync_pool/items/ and /async_pool/items/ that raced against ab's socket timeout. These 500s were silently included in the averages.

Fixes (harness + infra)

  • app_files/Dockerfile, app_files/gconf.py: socket-mode binds fixed, backlog=1024, timeout=300 — nginx 502s gone.
  • bin/run_test.sh: propagate container exit code (fail CI on error).
  • test_files/compare_container_performance.py: per-test socket_timeout / wall_timeout (600s/1200s only for the nginx + keepalive configs) and a per-test non-2xx tolerance guard (pool tests tolerate the known pool-exhaustion 500s, everything else stays strict).
  • test_files/test_nginx_port_vs_socket.py, test_files/test_keepalive.py, test_files/test_connection_pool.py: conflicting sockets/timeouts removed; pool tests annotated with the tolerance.

Doc rewrites (validated numbers replace the old ones)

Page Old (unreliable) New (validated)
nginx_port_socket flat ~7k rps via socket, 2.5-3.8x faster +9-19% on small payloads, +/-3% on 1MB
keepalive +1.79% sync / -2.3% async "regression" +8.98% sync, +12.30% async, ~0 on 1MB
workers_and_threads 446 rps w1 on 1MB, 6-7x gains, 3300-3600 rps ceiling 2 workers is the sweet spot; +50% small, +90% 1MB; threads flat
index +297%, ~4x combined +116-144% (2.2-2.4x) combined

Validation

  • Run 36235373898: all jobs green, every step success, no unexpected non-2xx responses.
  • Artifacts verified by parsing the HTML report blobs; every table in the docs matches the CI logs/artifacts.

Note: branch protection requires 3.10 context checks that no longer exist (the matrix is 3.14) — an admin override is required to merge.

… breaks unix socket binds)

GUNICORN_CMD_ARGS set --reuse-port for every container. On unix socket
binds gunicorn 26.0.0 drops the bind address and workers fail to
connect (ENOTSUP), so nginx proxied 502 for all requests. Socket-mode
(nginx_port_vs_socket, nginx_socker_keepalive) numbers were bogus.
Enable reuse_port only for TCP binds via gconf.py. Also propagate
pytest's exit code from run_test.sh so failed measurement steps no
longer silently pass CI.
…load

- raise gunicorn backlog 64->1024 (accept queue overflowed on the unix
  socket under -c100, nginx saw connect() EAGAIN and returned 502)
- raise gunicorn timeout 60->300 (workers were killed mid-request on the
  heavily CPU-bound 1MB async path, resetting connections into 502s)
- allow up to 1% transient non-2xx responses in the test guard, keeping
  the strict 307-trailing-slash check as a hard error
- raise ab socket timeout -s 60->600: under -c100 the upstream clients
  waited >60s behind a single serialized 1MB worker, ab dropped them
  mid-response, nginx aborted the upstream write and the worker wasted
  work on dead sockets - a self-reinforcing slowdown that left the
  container effectively wedged (0 responses, rps 0)
- raise the subprocess wall-cap 100s->1200s so slow-but-valid runs
  finish instead of being killed and discarded
- fail loudly with a clear message when baseline RPS is 0 (was an
  opaque ZeroDivisionError)
-scope -s 600 / subprocess cap 1200 to nginx benchmark configs only
-default stays -s 60 / cap 100, preserving pool/keepalive semantics
pool=2/40 containers 500 from httpx pool-timeout (60s) under 100 clients;
small-pool tests are expected to shed errors. Default stays strict 1%.
@github-actions
github-actions Bot enabled auto-merge September 27, 2026 00:01
@github-actions
github-actions Bot merged commit d68f496 into main Sep 27, 2026
15 of 18 checks passed
@github-actions
github-actions Bot deleted the fix/ci-measurements-validated-docs branch September 27, 2026 08:39
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant