Skip to content

perf: retune serving yield batch to 64 - #6

Merged
colmugx merged 2 commits into
mainfrom
perf/retune-yield-batch-64
Sep 24, 2026
Merged

colmugx merged 2 commits into
mainfrom
perf/retune-yield-batch-64

Conversation

@colmugx

@colmugx colmugx commented Sep 24, 2026

Copy link
Copy Markdown
Owner

Why

The current 32-message serving batch leaves measurable scheduler overhead on sustained backlogs.

A same-runner sweep of 32/64/128/256/512 showed that 64 improves tell throughput materially while keeping the existing cancellation/lifecycle tests intact. Batch 128 was rejected because it caused T10 (asker cancellation behind backlog) to complete the ask before cancellation was observed.

Measured result

On AMD EPYC 7763 with the same MoonBit toolchain:

  • tell median: 12.706 M/s -> 14.785 M/s (+16.4%)
  • handler-cost median: 516.7k/s -> 518.0k/s (flat)
  • ask-idle P99: 7 us -> 7 us
  • ask-loaded P99: 10 us -> 9 us
  • hotcold-16 P99: 125 us -> 217 us

What changed

  • normal serving batch: 32 -> 64
  • stranded-message cleanup batch remains 32
  • public API unchanged

The cleanup cadence is split deliberately so throughput tuning does not make abnormal cleanup less responsive.

Why: a same-runner sweep showed batch 64 improves tell throughput by about 16% over 32, while batch 128 violates the ask-cancellation timing contract.

What: use 64 for normal serving and keep stranded-message cleanup at 32.

Result: higher steady-backlog throughput without changing public APIs or weakening lifecycle tests.
@colmugx
colmugx merged commit 34a6352 into main Sep 24, 2026
1 check passed
@colmugx
colmugx deleted the perf/retune-yield-batch-64 branch September 24, 2026 04:45
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant