Area: CI / tests — the Integration tests job
On a PR, CI went from 3–4 minutes to about 9 when the distributed-backends work landed. The whole difference is the Integration tests job, which is now the critical path. Every other job is unchanged.
Measured
GitHub Actions, Integration tests job:
|
Before (a PR run on 2026-09-29) |
#624's last run (2026-09-30) |
tests/integration |
2m45s, 453 tests |
7m04s, 493 tests |
internal/mq integration run (a second gotestsum call, after the first) |
— |
57s |
internal/mq/natsspike |
— |
18s |
| Job total |
~3m |
8m44s |
| Whole CI run |
3–4m |
9m |
Over August and most of September, successful PR runs of ci.yml averaged 3.7–4.2 minutes (438 runs). Nothing else drifted.
Local: go test -tags integration -race -count=1 -json ./tests/integration/ on main @ 6fa9723 took 343s. The top-level tests add up to 313s, so the package runs almost serially. Only 5 of its files call t.Parallel(). The slowest tests:
| s |
test |
from |
| 62.1 |
TestRoles_SeparateProcesses |
#622/#624 |
| 39.9 |
TestShardClaims_TableOrderAcrossHandoverAndStop |
#624 |
| 29.6 |
TestShardClaims_BlockedOwnerKeepsItsShard |
#624 |
| 28.7 |
TestIngest_ClickHouseOutage_RetriedNotDeadLettered |
#619 |
| 21.4 |
TestShardClaims_StuckShardDoesNotStallTheOthers |
#624 |
| 20.4 |
TestShardClaims_CrashWithRowsInFlight |
#624 |
| 15.1 |
TestBootResilience_StickyHealthVsConditionalReady |
older |
| 11.5 |
TestNATSBackend_EndToEnd |
#624 |
| 10.7 |
TestQuery_MutationsReturnEmptyArray |
older |
| 7.3 + 6.5 |
TestSharedCache_* |
#614 |
| 6.5 |
TestShardClaims_ScaleOneThreeTwo |
#624 |
| 5.7 |
TestDynamoDBDedupe_TwoInstancesShareSeenIDs |
#625 |
The tests from the distributed work add up to about 215s, which accounts for the growth.
Why they are slow (inferred from the tests' code)
Most of that time is spent waiting on real clocks, not doing work:
- Lease and pin timers.
TestShardClaims_BlockedOwnerKeepsItsShard blocks an insert for 13s to outlast the 10s pinned TTL. TestShardClaims_TableOrderAcrossHandoverAndStop makes the old owner's inserts take 6s each. The crash test waits out a pin TTL. TestRoles_SeparateProcesses waits for real membership leases (15s unrenewed) across its handover steps, in sequence.
- Serial execution. None of these tests calls
t.Parallel(), although each starts its own NATS server, and most start their own processes and fake ClickHouse.
- Serial
gotestsum calls. The test-integration Makefile target runs internal/mq's integration tests as a second call, after tests/integration finishes.
Fix directions, roughly cheapest first
- Run the pieces concurrently. Split
Integration tests into matrix jobs: tests/integration on its own; internal/mq integration plus internal/cache plus natsspike. Or shard tests/integration by -run pattern. .github/workflows/README.md has a deferred sharding design with a working implementation (see "Deferred optimizations"). Coverage merges already happen in the Coverage job from artifacts. This alone should bring the job to about the slowest shard.
t.Parallel() on the self-contained slow tests: TestShardClaims_*, TestRoles_SeparateProcesses, TestNATSBackend_EndToEnd, TestSharedCache_*. Check each first for shared state: fixed ports, the shared ClickHouse container's tables (use per-test table names), and TestMain-built binaries. The -race CPU cost on a 4-core runner caps how much this buys, so measure it.
- Shorten the clocks under test. Where a test waits on a production default (pinned TTL 10s, lease 15s, ack_wait 60s), pass the smallest value the code accepts through the config the test already builds, and scale the sleeps to match. The pinned-TTL floor of 10s is a check at boot; a test-only override would need care so it cannot reach production config.
- Decide whether
internal/mq/natsspike belongs in CI at all. It is an experiment package: 18s per run.
- Local
make ci pays the same cost, and it is the pre-push gate, so this slows every push too, not just the PR check.
Done when
Related: #624, #613, #692.
Area: CI / tests — the Integration tests job
On a PR, CI went from 3–4 minutes to about 9 when the distributed-backends work landed. The whole difference is the
Integration testsjob, which is now the critical path. Every other job is unchanged.Measured
GitHub Actions,
Integration testsjob:tests/integrationinternal/mqintegration run (a secondgotestsumcall, after the first)internal/mq/natsspikeOver August and most of September, successful PR runs of
ci.ymlaveraged 3.7–4.2 minutes (438 runs). Nothing else drifted.Local:
go test -tags integration -race -count=1 -json ./tests/integration/onmain@ 6fa9723 took 343s. The top-level tests add up to 313s, so the package runs almost serially. Only 5 of its files callt.Parallel(). The slowest tests:TestRoles_SeparateProcessesTestShardClaims_TableOrderAcrossHandoverAndStopTestShardClaims_BlockedOwnerKeepsItsShardTestIngest_ClickHouseOutage_RetriedNotDeadLetteredTestShardClaims_StuckShardDoesNotStallTheOthersTestShardClaims_CrashWithRowsInFlightTestBootResilience_StickyHealthVsConditionalReadyTestNATSBackend_EndToEndTestQuery_MutationsReturnEmptyArrayTestSharedCache_*TestShardClaims_ScaleOneThreeTwoTestDynamoDBDedupe_TwoInstancesShareSeenIDsThe tests from the distributed work add up to about 215s, which accounts for the growth.
Why they are slow (inferred from the tests' code)
Most of that time is spent waiting on real clocks, not doing work:
TestShardClaims_BlockedOwnerKeepsItsShardblocks an insert for 13s to outlast the 10s pinned TTL.TestShardClaims_TableOrderAcrossHandoverAndStopmakes the old owner's inserts take 6s each. The crash test waits out a pin TTL.TestRoles_SeparateProcesseswaits for real membership leases (15s unrenewed) across its handover steps, in sequence.t.Parallel(), although each starts its own NATS server, and most start their own processes and fake ClickHouse.gotestsumcalls. Thetest-integrationMakefile target runsinternal/mq's integration tests as a second call, aftertests/integrationfinishes.Fix directions, roughly cheapest first
Integration testsinto matrix jobs:tests/integrationon its own;internal/mqintegration plusinternal/cacheplusnatsspike. Or shardtests/integrationby-runpattern..github/workflows/README.mdhas a deferred sharding design with a working implementation (see "Deferred optimizations"). Coverage merges already happen in theCoveragejob from artifacts. This alone should bring the job to about the slowest shard.t.Parallel()on the self-contained slow tests:TestShardClaims_*,TestRoles_SeparateProcesses,TestNATSBackend_EndToEnd,TestSharedCache_*. Check each first for shared state: fixed ports, the shared ClickHouse container's tables (use per-test table names), andTestMain-built binaries. The-raceCPU cost on a 4-core runner caps how much this buys, so measure it.internal/mq/natsspikebelongs in CI at all. It is an experiment package: 18s per run.make cipays the same cost, and it is the pre-push gate, so this slows every push too, not just the PR check.Done when
gh run list -w ci.ymlbefore and after.covgate).internal/mq's integration tests are chosen by name pattern. If the jobs are restructured, fix that selection at the same time.Related: #624, #613, #692.