Skip to content

Add uio-ws (Go) engine entry - #1522

Open
limpo1989 wants to merge 4 commits into
MDA2AV:mainfrom
limpo1989:main
Open

limpo1989 wants to merge 4 commits into
MDA2AV:mainfrom
limpo1989:main

Conversation

@limpo1989

Copy link
Copy Markdown

Description

uio-ws — Go engine entry

This adds a WebSocket echo entry for uws, the WebSocket implementation on uio (github.com/urpc/uio, v1.5.3).

uio is a Linux/BSD network engine for Go built directly on epoll (EPOLLET) and kqueue (EV_CLEAR): one edge-triggered event loop per CPU, each connection driven by its own serialized I/O task — no goroutine per connection, no reactor-to-worker handoff. uws implements RFC 6455 framing and message rules plus RFC 7692 permessage-deflate on top of it, and passes the Autobahn test suite.

What the echo path looks like:

  • The HTTP/1.1 upgrade handshake and the frame parser run inside the connection's own I/O task, so a message is parsed straight from the read buffer its task filled and echoed from that same task.
  • A read round's replies are corked and coalesced: a burst of frames costs one sendmsg (writev-style iovecs, stack-backed) instead of a syscall per reply — which is what echo-ws-pipeline exercises.
  • A send the socket refuses is handed to the next write turn instead of pausing the read, so a slow peer never stalls the reader; outbound backpressure is bounded by memory limits rather than blocking.
  • Short-lived connections — open, a few frames, close — are a first-class path (v1.5.3 fixed reading a FIN queued behind the last bytes under edge triggering, exactly the pattern echo-ws-limited drives).

Entry notes:

  • type: engine — uio is a transport engine applications are not written against; happy to move if that reads wrong.
  • Subscribes to echo-ws, echo-ws-pipeline, echo-ws-limited.
  • Configuration is uws defaults; the only explicit setting is Pollers: runtime.NumCPU() (uio's default is 4 loops).
  • scripts/validate.sh uio-ws passes 7/7 WebSocket checks; the entry is a 60-line echo handler, no benchmarking tricks.
  • Maintainer: @limpo1989

PR Commands — comment on this PR to trigger (requires collaborator approval):

Command Description
/benchmark -f <framework> Run every test the framework subscribes to
/benchmark -f <framework> -t <test> Run one test only
/benchmark -f <framework> --save Run and save results (updates the leaderboard on merge)
/benchmark -f <framework> -t <test> --save Run one test and save results
/benchmark -f <framework> --compare <other> Measure the deltas against another framework instead of this one
/benchmark-multiple -f <fw1>,<fw2>,... Benchmark several frameworks in one run — takes -t and --save too; saved results land in a single commit
/benchmark-multiple --save No -f needed: benchmark and save every framework the PR touches
/benchmark-test -t <test> Benchmark all enabled frameworks subscribed to <test> and save the results

For /benchmark, always specify -f <framework>; the flags combine in any order. Results come back as a comment with a per-profile table of RPS, p99, CPU and memory — one table per framework on multi runs. A new benchmark comment while a run is in flight queues behind it (one deep) instead of cancelling it. For multi-framework PRs (dependency bumps, same-language refactors) prefer /benchmark-multiple, which runs everything in a single job and commits all saved results together, so no run overwrites another. --compare works on single-framework runs only.

What the deltas are measured against. By default, this framework's own results published on main - answering "did this change help?". When you are tuning a variant or a successor entry, --compare re-bases them on another entry instead:

/benchmark -f genhttp --compare genhttp-kestrel

The reply states which baseline it used, and profiles the other framework does not run show n/a rather than a delta.


Run benchmarks locally

You can validate and benchmark your framework locally with the lite script — no CPU pinning, fixed connection counts, all load generators run in Docker.

./scripts/validate.sh <framework>
./scripts/benchmark-lite.sh <framework> baseline
./scripts/benchmark-lite.sh --load-threads 4 <framework>

Requirements: Docker Engine on Linux. Load generators (gcannon, h2load, h2load-h3, wrk) are built as self-contained Docker images on first run.

uws WebSocket echo on uio: the RFC 6455 handshake and frame parser
run in the connection's own I/O task on uio's edge-triggered event
loops, one per CPU, so each message is parsed from the buffer its
task read into and echoed from that task with no goroutine per
connection. A read round's replies leave coalesced in one sendmsg,
and a send the socket refuses is finished by the next write turn
without pausing the read. Everything else is default configuration;
uio v1.5.3.

Subscribes to echo-ws, echo-ws-pipeline and echo-ws-limited.
Validated with scripts/validate.sh (7/7 WebSocket checks).
@MDA2AV

MDA2AV commented Oct 4, 2026

Copy link
Copy Markdown
Owner

/benchmark --save

@github-actions

github-actions Bot commented Oct 4, 2026

Copy link
Copy Markdown
Contributor

👋 Benchmark request received. A collaborator will review and approve the run.

@github-actions

github-actions Bot commented Oct 4, 2026

Copy link
Copy Markdown
Contributor

Benchmark Results

Framework: uio-ws | Test: all tests

Test Conn RPS CPU Mem Δ RPS Δ Mem
echo-ws 512 1,866,244 3276.6% 51MiB ~0% ~0%
echo-ws 4096 1,957,550 3307.9% 77MiB ~0% ~0%
echo-ws 16384 1,519,552 3035.5% 165MiB ~0% ~0%
echo-ws-pipeline 512 29,343,251 3681.4% 51MiB ~0% ~0%
echo-ws-pipeline 4096 30,318,937 3765.6% 80MiB ~0% ~0%
echo-ws-pipeline 16384 23,882,179 3121.0% 171MiB ~0% ~0%
echo-ws-limited 512 410,035 1465.1% 64MiB ~0% ~0%
echo-ws-limited 4096 521,759 1441.5% 94MiB ~0% ~0%
Full log
  Bandwidth:  159.13MB/s
  WS upgrades: 16384
  WS frames:   118930832
  Latency samples: 118930832 / 118930832 responses (100.0%)
[info] CPU 3390.0% | Mem 166MiB

[run 3/3]
gcannon v0.5.3 [WS]
  Target:    localhost:8080/ws
  Threads:   64
  Conns:     16384 (256/thread)
  Pipeline:  16
  Req/conn:  unlimited (keep-alive)
  Expected:  200
  Duration:  5s


  Thread Stats   Avg      p50      p90      p99    p99.9
    Latency   10.64ms   10.60ms   11.50ms   13.50ms   17.90ms

  119673040 frames sent in 5.00s, 119410896 frames received
  Throughput: 23.87M req/s
  Bandwidth:  159.78MB/s
  WS upgrades: 16384
  WS frames:   119410896
  Latency samples: 119410896 / 119410896 responses (100.0%)
[info] CPU 3121.0% | Mem 171MiB

=== Best: 23882179 req/s (CPU: 3121.0%, Mem: 171MiB) ===
[info] saved results/echo-ws-pipeline/16384/uio-ws.json
httparena-bench-uio-ws
httparena-bench-uio-ws

==============================================
=== uio-ws / echo-ws-limited / 512c (tool=gcannon) ===
==============================================
[info] ws-only framework — skipping HTTP probe (sleep 2s for startup)

[run 1/3]
gcannon v0.5.3 [WS]
  Target:    localhost:8080/ws
  Threads:   64
  Conns:     512 (8/thread)
  Pipeline:  1
  Req/conn:  10
  Expected:  200
  Duration:  5s


  Thread Stats   Avg      p50      p90      p99    p99.9
    Latency    920us    683us   1.82ms   3.37ms   5.04ms

  2050112 frames sent in 5.00s, 2050176 frames received
  Throughput: 409.94K req/s
  Bandwidth:  7.78MB/s
  WS upgrades: 207366
  WS frames:   2050176
  Latency samples: 2050176 / 2050176 responses (100.0%)
  Reconnects: 205034
[info] CPU 1465.1% | Mem 64MiB

[run 2/3]
gcannon v0.5.3 [WS]
  Target:    localhost:8080/ws
  Threads:   64
  Conns:     512 (8/thread)
  Pipeline:  1
  Req/conn:  10
  Expected:  200
  Duration:  5s


  Thread Stats   Avg      p50      p90      p99    p99.9
    Latency    918us    682us   1.80ms   3.46ms   5.32ms

  2038564 frames sent in 5.00s, 2038611 frames received
  Throughput: 407.61K req/s
  Bandwidth:  7.74MB/s
  WS upgrades: 207698
  WS frames:   2038611
  Latency samples: 2038605 / 2038611 responses (100.0%)
  Reconnects: 203890
[info] CPU 1484.3% | Mem 65MiB

[run 3/3]
gcannon v0.5.3 [WS]
  Target:    localhost:8080/ws
  Threads:   64
  Conns:     512 (8/thread)
  Pipeline:  1
  Req/conn:  10
  Expected:  200
  Duration:  5s


  Thread Stats   Avg      p50      p90      p99    p99.9
    Latency    927us    687us   1.84ms   3.37ms   5.19ms

  2046785 frames sent in 5.00s, 2046793 frames received
  Throughput: 409.23K req/s
  Bandwidth:  7.77MB/s
  WS upgrades: 208848
  WS frames:   2046793
  Latency samples: 2046780 / 2046793 responses (100.0%)
  Reconnects: 204641
[info] CPU 1487.2% | Mem 65MiB

=== Best: 410035 req/s (CPU: 1465.1%, Mem: 64MiB) ===
[info] saved results/echo-ws-limited/512/uio-ws.json
httparena-bench-uio-ws
httparena-bench-uio-ws

==============================================
=== uio-ws / echo-ws-limited / 4096c (tool=gcannon) ===
==============================================
[info] ws-only framework — skipping HTTP probe (sleep 2s for startup)

[run 1/3]
gcannon v0.5.3 [WS]
  Target:    localhost:8080/ws
  Threads:   64
  Conns:     4096 (64/thread)
  Pipeline:  1
  Req/conn:  10
  Expected:  200
  Duration:  5s


  Thread Stats   Avg      p50      p90      p99    p99.9
    Latency   1.70ms   1.30ms   3.30ms   6.04ms   8.71ms

  2590426 frames sent in 5.00s, 2589278 frames received
  Throughput: 517.73K req/s
  Bandwidth:  9.84MB/s
  WS upgrades: 260627
  WS frames:   2589278
  Latency samples: 2589278 / 2589278 responses (100.0%)
  Reconnects: 258420
[info] CPU 1430.4% | Mem 94MiB

[run 2/3]
gcannon v0.5.3 [WS]
  Target:    localhost:8080/ws
  Threads:   64
  Conns:     4096 (64/thread)
  Pipeline:  1
  Req/conn:  10
  Expected:  200
  Duration:  5s


  Thread Stats   Avg      p50      p90      p99    p99.9
    Latency   1.71ms   1.31ms   3.30ms   6.30ms   9.16ms

  2602159 frames sent in 5.00s, 2602563 frames received
  Throughput: 520.39K req/s
  Bandwidth:  9.87MB/s
  WS upgrades: 262936
  WS frames:   2602563
  Latency samples: 2602549 / 2602563 responses (100.0%)
  Reconnects: 260352
[info] CPU 1491.9% | Mem 94MiB

[run 3/3]
gcannon v0.5.3 [WS]
  Target:    localhost:8080/ws
  Threads:   64
  Conns:     4096 (64/thread)
  Pipeline:  1
  Req/conn:  10
  Expected:  200
  Duration:  5s


  Thread Stats   Avg      p50      p90      p99    p99.9
    Latency   1.71ms   1.30ms   3.36ms   6.33ms   9.61ms

  2609213 frames sent in 5.00s, 2608799 frames received
  Throughput: 521.62K req/s
  Bandwidth:  9.90MB/s
  WS upgrades: 262972
  WS frames:   2608799
  Latency samples: 2608798 / 2608799 responses (100.0%)
  Reconnects: 260628
[info] CPU 1441.5% | Mem 94MiB

=== Best: 521759 req/s (CPU: 1441.5%, Mem: 94MiB) ===
[info] saved results/echo-ws-limited/4096/uio-ws.json
httparena-bench-uio-ws
httparena-bench-uio-ws
[info] skip: uio-ws does not subscribe to latency-1m
[info] skip: uio-ws does not subscribe to latency-10k
[info] skip: uio-ws does not subscribe to latency-500k-8cpu
[info] skip: uio-ws does not subscribe to async
[info] rebuilding site/data/*.json
[updated] /home/diogo/actions-runner/_work/HttpArena/HttpArena/site/data/frameworks.json
[updated] /home/diogo/actions-runner/_work/HttpArena/HttpArena/site/data/results/uio-ws.json - 8 new, 8 total
[updated] /home/diogo/actions-runner/_work/HttpArena/HttpArena/site/data/current.json
[info] done
[info] restoring loopback MTU to 65536

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants