TestDedupeDynamo_Conformance/a_failed_reserve_leaves_nothing_claimed failed once in a local integration run (measured: 1 failure in 3 full runs of tests/integration on 2026-09-25, on the #645 tree; the test and internal/dedupe/dynamodb.go are unchanged from main). Key a was still InFlight after the failed Reserve:
dedupetest.go:297: expected: []dedupe.Status{0x1, 0x1, 0x1}
actual : []dedupe.Status{0x3, 0x1, 0x1}
So a pending item for a survived the rollback, and it holds the id until the lease lapses.
Likely cause (inferred, not verified). dynamoStore.Reserve sends its puts under the errgroup's gctx, and the first failure cancels that context. A put that was already on the wire when the context was cancelled returns context.Canceled to the client, but the server can still apply it. The rollback then sends a DeleteItem conditional on the put's token, on another connection. If dynamodb-local handles that delete before the cancelled put, the delete fails its condition and is treated as done, and the put then lands. The unit test TestDynamo_FailedMultiKeyReserve cannot see this because its fake returns at once when the context is cancelled.
The same thing can happen against real DynamoDB. The cost is an id that answers 503 InFlight for one dedupe.lease (30s by default), not a lost or duplicated row.
Possible fixes, for whoever owns internal/dedupe: let puts that were already sent run to completion under the caller's ctx, and use gctx only to stop unsent puts. This changes what TestDynamo_FailedMultiKeyReserve pins, so its fake would need a bound. Or, after the rollback, re-read each rolled-back key and delete any pending item that carries our token.
Found while fixing #645's integration timeout. That fix does not touch this code.
TestDedupeDynamo_Conformance/a_failed_reserve_leaves_nothing_claimedfailed once in a local integration run (measured: 1 failure in 3 full runs oftests/integrationon 2026-09-25, on the #645 tree; the test andinternal/dedupe/dynamodb.goare unchanged from main). Keyawas stillInFlightafter the failedReserve:So a pending item for
asurvived the rollback, and it holds the id until the lease lapses.Likely cause (inferred, not verified).
dynamoStore.Reservesends its puts under the errgroup'sgctx, and the first failure cancels that context. A put that was already on the wire when the context was cancelled returnscontext.Canceledto the client, but the server can still apply it. The rollback then sends aDeleteItemconditional on the put's token, on another connection. If dynamodb-local handles that delete before the cancelled put, the delete fails its condition and is treated as done, and the put then lands. The unit testTestDynamo_FailedMultiKeyReservecannot see this because its fake returns at once when the context is cancelled.The same thing can happen against real DynamoDB. The cost is an id that answers
503 InFlightfor onededupe.lease(30s by default), not a lost or duplicated row.Possible fixes, for whoever owns
internal/dedupe: let puts that were already sent run to completion under the caller'sctx, and usegctxonly to stop unsent puts. This changes whatTestDynamo_FailedMultiKeyReservepins, so its fake would need a bound. Or, after the rollback, re-read each rolled-back key and delete any pending item that carries our token.Found while fixing #645's integration timeout. That fix does not touch this code.