ALCF CI: trigger a fresh GitLab pipeline on re-runs after an unsuccessful pipeline - #648
Merged
Merged
Conversation
A mirror sync does not create a new pipeline for a SHA that GitLab already has, so re-running the GitHub check re-adopted the old pipeline's result, e.g. one whose jobs failed with stuck_pending_no_matching_runners during an Aurora runner outage. On re-runs, create a fresh pipeline via the trigger token unless one is still active for the SHA.
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #648 +/- ##
==========================================
+ Coverage 78.95% 80.72% +1.76%
==========================================
Files 56 56
Lines 4087 4088 +1
==========================================
+ Hits 3227 3300 +73
+ Misses 860 788 -72 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
Trigger a fresh pipeline on re-runs only when the newest pipeline for the head SHA failed, was canceled or was skipped. A missing, active or successful pipeline, or a failed lookup, is left to the wait step, so a transient API failure can no longer create a duplicate pipeline. Replace the mirror wait loop with a single branch-tip check: a finished pipeline for the SHA means GitLab already has the commit, and without one the mirror sync creates the pipeline itself. Also verify that the triggered pipeline runs the head SHA, and pass the ref with --form-string so curl does not interpret it.
michel2323
marked this pull request as ready for review
September 29, 2026 15:04
michel2323
enabled auto-merge (squash)
September 29, 2026 16:19
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Re-running the ALCF CI check could never recover from a GitLab pipeline that failed for infrastructure reasons. For example, in run 36428123880 both
build:supportjobs failed withstuck_pending_no_matching_runnersduring an Aurora runner outage. The mirror sync doesn't create a new pipeline for a SHA that GitLab already has, so the wait step kept finding the dead pipeline 66691 and failing within seconds.On re-runs (
run_attempt > 1), this PR creates a fresh pipeline through the trigger token when the newest pipeline for the SHA failed, was canceled or was skipped. In every other case (no pipeline yet, one still active, one that succeeded, or a failed lookup) the check adopts whatever the wait step finds, as before. Retrying the old pipeline instead isn't possible, because the retry would be attributed to the bot user and the Jacamar runners refuse those jobs.The trigger runs the branch's current tip, so the step fails without triggering if the GitLab branch tip is not the head SHA, and it verifies that the triggered pipeline runs the head SHA.
Testing: let this PR's pipeline end failed or canceled (for example by canceling it in GitLab), re-run the check, and confirm that the log shows
triggered fresh pipeline <id>and the check adopts that new pipeline. Re-running while the pipeline is active or after it succeeded should lognothing to re-trigger.