valkey - test: Run cluster tests on their own Valkey cluster - #2180
Merged
Merged
Conversation
`pnpm test:ci` runs every package's tests at once, and @keyv/redis and @keyv/valkey both ran their cluster tests against the Redis cluster on 7001-7003. The Redis tests clear that cluster with FLUSHDB before each test, so Valkey test keys vanished mid-test. On #2179 this failed the node-26 job; retries had already masked three other hits in that run. Add a three-node Valkey 9.1.0 cluster on 7101-7103, started and stopped with the other test services, and point the Valkey cluster tests at it. They now also run against Valkey rather than Redis. Drop the per-test retries that were hiding the race. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01X5LJCvt3pkR7FtfyAdzg5x
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #2180 +/- ##
=========================================
Coverage 100.00% 100.00%
=========================================
Files 56 56
Lines 5790 5790
Branches 997 989 -8
=========================================
Hits 5790 5790 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
2 tasks done
jaredwray
added a commit
that referenced
this pull request
Oct 3, 2026
Seven tests in six adapters set a value that expires in 100 ms, then read it back expecting it to still be there. When set() and get() together take longer than 100 ms under CI load, the read finds the value expired. The postgres copy needed a retry in the #2180 branch run. - postgres, mysql, sqlite and cloudflare-kv decide expiry with Date.now(), so these tests now freeze Date and move it past the TTL instead of sleeping. They can't race the clock, and run 200 ms faster. - redis and valkey expire keys on the server's clock, so those tests get a 1 s window instead of 100 ms. The postgres and mysql copies also turn off Keyv's own expiry check, as the sqlite copy already did. With it on, Keyv deleted the expired value itself, so the tests passed even when the adapter left the expires column empty, the bug they are meant to catch. Claude-Session: https://claude.ai/code/session_01X5LJCvt3pkR7FtfyAdzg5x Co-authored-by: Claude <noreply@anthropic.com>
jaredwray
pushed a commit
that referenced
this pull request
Oct 3, 2026
Nine storage packages retried every failed test twice. That hid real problems: the valkey cluster tests were losing keys to @keyv/redis's FLUSHDB (#2180) and several tests raced a 100 ms TTL (#2182), and both only surfaced when a test failed three times in a row. With those fixed, a test that needs a retry to pass should fail CI so it gets fixed. Remove `retry: 2` from the vitest configs of cloudflare-kv, dynamo, etcd, memcache, mongo, mysql, postgres, redis and valkey. The Cloudflare KV live config keeps its retries: it calls the real Cloudflare API over the internet. AGENTS.md now says not to add retries to get CI green. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01X5LJCvt3pkR7FtfyAdzg5x
2 tasks done
jaredwray
added a commit
that referenced
this pull request
Oct 3, 2026
…2184) Nine storage packages retried every failed test twice. That hid real problems: the valkey cluster tests were losing keys to @keyv/redis's FLUSHDB (#2180) and several tests raced a 100 ms TTL (#2182), and both only surfaced when a test failed three times in a row. With those fixed, a test that needs a retry to pass should fail CI so it gets fixed. Remove `retry: 2` from the vitest configs of cloudflare-kv, dynamo, etcd, memcache, mongo, mysql, postgres, redis and valkey. The Cloudflare KV live config keeps its retries: it calls the real Cloudflare API over the internet. AGENTS.md now says not to add retries to get CI green. Claude-Session: https://claude.ai/code/session_01X5LJCvt3pkR7FtfyAdzg5x Co-authored-by: Claude <noreply@anthropic.com>
8 of 10 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Please check if the PR fulfills these requirements
@keyv/valkey's coverage is the same as onmain.What kind of change does this PR introduce? (Bug fix, feature, docs update, ...)
Test infrastructure fix (a CI race).
Problem
pnpm test:ciruns every package's tests at the same time.@keyv/redisand@keyv/valkeyboth ran their cluster tests against the one Redis cluster on ports 7001-7003. In each of the seven CI job logs I checked, the two cluster test files finished within a second of each other. Before each of its 23 tests, the Redis suite clears the whole cluster (clear()withnoNamespaceAffectsAll, which runsFLUSHDBon every master). Keys the Valkey tests had just written disappeared mid-test.On #2179 this failed the node-26 job.
should track keys with useSets without CROSSSLOT errorsfailed all four attempts atexpect(await store.delete(keys[0])).toBe(true): the key set one line earlier was already gone. The same run retried three more Valkey cluster tests (deleteMany,clear,iterator) that then passed. The per-test{ retry: 3 }was hiding most of the hits.Local reproduction: a three-node cluster on 7001-7003, with
@keyv/redis's ownclear()running in a loop (what itsbeforeEachdoes) while the Valkey cluster tests run with retries off. 3 of 5 runs failed.Changes
scripts/docker-compose-valkey-cluster.yaml(new):keyv_valkeyservice;valkey-cli --cluster create.scripts/test-services-start.shandtest-services-stop.sh: start and stop the new cluster with the other services, on x86 and ARM. Every workflow that needs services (tests, codecov, bun-test, release) goes throughpnpm test:services:start.storage/valkey/test/cluster.test.ts:{ retry: 3 }that was hiding the race. The package-wideretry: 2invitest.config.tsstays.AGENTS.mdandCONTRIBUTING.md:No public API or adapter behavior changes, so the migration guide and skill are unchanged.
For local runs: after pulling this, run
pnpm test:services:startagain to start the new cluster.docker compose up -dadds it without touching the running services.Verification
clear()looping on 7001-7003 while the Valkey tests ran on 7101-7103. 0 of 10 runs failed, with 1,050 flushes during the runs.redis-server7.0 standing in for Valkey. CI runs the real Valkey image.pnpm test:ciinstorage/valkey: 158/158 pass, lint is clean, and coverage matchesmain's CI.docker compose config: parses the merged files for both the x86 and ARM service sets.🤖 Generated with Claude Code
https://claude.ai/code/session_01X5LJCvt3pkR7FtfyAdzg5x
Generated by Claude Code