Re-queue behind a dead BullMQ job, not just an absent one - #286
Merged
Merged
Conversation
The sweep asked whether a BullMQ job existed for a stranded row and skipped the row if one did. Existence is the wrong question: BullMQ retains a job after it finishes, so a failed or completed job is still findable by id and is emphatically not scheduled. Every row whose attempt had failed sat queued forever behind the corpse of the attempt that failed it — which is exactly the stranding this sweep was written to undo. It cost a working deploy. With the GIF frame-pattern bug fixed and shipped and the row reset to queued, nothing rendered for five minutes and the sweep logged nothing, because it was skipping the row on every pass. A job in a live state is still left alone. A dead one is removed so its id is free, and the row is scheduled again. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
ThreatCrush Security Scan48 finding(s) HIGH/CRITICAL: 2 | MEDIUM: 31 | LOW: 15
Snippets are redacted; ThreatCrush never prints matched credential material. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The sweep asked whether a BullMQ job existed for a stranded row, and skipped the row if one did.
Existence is the wrong question. BullMQ retains a job after it finishes, so a failed or completed job is still findable by id — and is emphatically not scheduled. Every row whose attempt had failed sat
queuedforever behind the corpse of the attempt that failed it, which is exactly the stranding this sweep was written to undo.It cost a working deploy
With the GIF frame-pattern bug fixed, shipped, and the row reset to
queued, nothing rendered for five minutes and the sweep logged nothing at all — because it was skipping the row on every pass. Three separate fixes were live and unreachable.The change
A job in a live state (
waiting,active,delayed) is still left alone. A dead one —failedorcompleted— is removed so its id is free, and the row is scheduled again.Verification
2565 passed / 1 failed repo-wide — the pre-existing
tracker-geofailure. Root and worker typechecks clean.This is the fourth bug in this chain, and every one of them was found by queueing a real render and reading the worker log rather than trusting a green pipeline. None of them could have been caught by CI: the frame-pattern mismatch needed the real capturer, the attempts cap needed a real failed row, and this one needs a real Redis with a retained job.