Conversation
…rs, at least 8 Assisted-by: Claude Code (claude-opus-5-5)
This was referenced Oct 4, 2026
11 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Motivation
The broker's Netty I/O threads (
numIOThreads, thepulsar-ioevent loops) default to2 * Runtime.getRuntime().availableProcessors(): 16 on an 8-processor host and 32 on a 16-processor one. In the Pulsar Performance Testing Framework's IoT telemetry scenarios, that many event loops cost CPU and throughput:epoll_waitmore often and batches less per read and write system call.ServerCnx.execute), and every bookie write to the BookKeeper client's channel. With many mostly idle loops, most of those handovers wake a sleeping loop with aneventfdwrite.Too few I/O threads hurt too: they can't keep up with dispatching to the consumers. So the default needs a floor.
Modifications
New default: half the available processors, but no fewer than 8 unless that exceeds twice the processors:
Math.min(2 * n, Math.max(8, n / 2))for n available processors.Up to 4 processors, the default doesn't change.
Code:
ServiceConfiguration.defaultNumIOThreads(int)computes the default from one read of the available processors. A test covers the default for 1 to 128 processors.Docs: the documentation in
ServiceConfiguration,broker.conf,standalone.confand the Terraform template states the formula, and names the symptom of too few I/O threads: the busiestpulsar-iothread stays busy and consumers' backlogs grow.Buffer share:
ServerCnxdividesmaxMessagePublishBufferSizeInMBbetween the I/O threads, so each thread's share grows with fewer threads; the total is unchanged.Unchanged: the proxy's
numIOThreads, the BookKeeper client'sbookkeeperClientNumIoThreadsand the HTTP server's threads. An explicitly configurednumIOThreadsstill wins.This changes a default for deployments with more than 4 available processors, so it needs a release note.
Measurements
The framework's IoT telemetry scenarios, in the test cluster's single broker:
iot-telemetry-max-rate: 500 gateways publish unbatched messages to one topic without a rate limit, with up to 100,000 in flight, and 20 consumers on one Key_Shared subscription receive them.iot-telemetry-high-rate: 30,000 msg/s to 30 topics, with 5 subscriptions of 10 consumers each.Setup:
numIOThreadsvalues.docker update --cpuset-cpus, and setting-XX:ActiveProcessorCountto match.Catch-up with fewer threads: with the gateways connecting concurrently, 6 and 4 I/O threads caught up faster still, at a higher tail latency for the application reading at the tail. The table has means of 2 runs.
Why the floor stays at 8: on smaller brokers, 4 aren't enough at the maximum rate (above).
Max rate 128 B on a broker with 8 hardware threads
Max rate 128 B on a broker with 4 hardware threads: why the default doesn't go below 8
Throughput at max rate 128 B on 16 hardware threads, numIOThreads 32 (A) and 8 (B), with the same axes
Extrapolated: above 17 processors, the new default is half the processors, which nothing measured. Up to 17, the measurements support the floor of 8.
Shared with other uses: the broker's I/O event loop group also serves the broker's internal client for replication and lookups. Protocol handlers with a dedicated worker group get
numIOThreadsthreads too.Not measured: TLS (handshakes and encryption run on the event loops), 128 KB entries, thousands of connections or topics, a reconnect storm, geo-replication, and protocol handlers with dedicated worker groups sized from
numIOThreads. The test host runs the clients and bookies beside the broker, so part of the gain on 16 hardware threads may come from less CPU contention, which a broker on a host of its own wouldn't have. The pinned runs model a broker on cores of its own.Verifying this change
This change added tests and can be verified as follows:
ServiceConfigurationDefaultsTest:numIOThreadsdefaults to it.Does this pull request potentially affect one of the following parts:
If the box was checked, please highlight the changes
numIOThreads: half the available processors, but no fewer than 8 unless that exceeds twice the processors, instead of twice the processors.This change was prepared with the assistance of Claude Code (claude-opus-5-5); I have reviewed and verified it.