Skip to content

Prioritize error requeues with less maintainer noise and clear resolved failures - #853

Open
majamassarini wants to merge 2 commits into
packit:mainfrom
majamassarini:fix/error-list-requeue-and-cleanup
Open

majamassarini wants to merge 2 commits into
packit:mainfrom
majamassarini:fix/error-list-requeue-and-cleanup

Conversation

@majamassarini

@majamassarini majamassarini commented Sep 28, 2026 •

Copy link
Copy Markdown
Member

Summary

  • Reset retry attempts and send all make requeue-error tasks through the priority _todo queue. By default, keep user_triggered=false to reduce acknowledgement and result comments for maintainers; USER_TRIGGERED=true opts into that feedback.
  • Mark requeued tasks so triage and reproducer process them even while the issue still has an errored label.
  • Remove matching entries from error_list when triage, rebase, backport, rebuild, or reproducer work succeeds. Match by issue, workflow queue, and target branch to preserve unrelated errors.
  • After a successful consolidated rebase or rebuild, clear matching errors for the primary issue and each distinct resolved sibling.
  • For reproducer work, also clear matching errors after finalized created, already exists, or not reproducible outcomes. Keep failures and deferred or retryable results in the error list.
  • Retain reproducer errors when adapting an existing test fails during its merge request update, even if test_already_exists remains true.
  • Keep error_list untouched when DRY_RUN=true, including on successful processing.

@qodo-for-packit

Copy link
Copy Markdown

PR Summary by Qodo

Fix error requeue routing and clear resolved workflow failures

🐞 Bug fix 🧪 Tests 🕐 40+ Minutes

Grey Divider

AI Description

• Requeue errors onto normal queues by default or priority queues when explicitly requested.
• Reset retry attempts and let requeued triage and reproducer tasks run despite existing error
 labels.
• Clear matching workflow failures after success without removing errors for other issues, queues,
 or branches.
Diagram

graph TD
  E[("Redis error list")] --> C["Requeue command"] --> Q["Normal or priority queue"] --> A["Workflow agents"] --> U["Success cleanup"]
  U -->|"remove matching failures"| E
Loading
High-Level Assessment

Keep success-triggered cleanup in each agent, with matching centralized in the shared helper: each agent knows when its domain-specific work has succeeded, while the helper consistently preserves unrelated failures. A generic task-loop cleanup hook would need additional workflow-specific success and branch information.

Files changed (11) +202 / -18

Enhancement (1) +5 / -0
models.pyIdentify tasks requeued from the error list +5/-0

Identify tasks requeued from the error list

• Task gains a default-false marker that lets triage and reproducer distinguish operator requeues from duplicate work without changing user-triggered behavior.

ymir/common/models.py

Bug fix (8) +110 / -18
MakefileExpose optional priority routing for error requeues +2/-2

Expose optional priority routing for error requeues

• The requeue target accepts USER_TRIGGERED=true and passes the corresponding flag to the script. Its usage message documents the option.

openshift/Makefile

requeue_error.pyRoute requeued tasks according to the requested mode +27/-16

Route requeued tasks according to the requested mode

• Requeues now default to the ordinary queue, with --user-triggered selecting its priority counterpart. Both modes reset attempts and mark the task as requeued from the error list.

openshift/scripts/requeue_error.py

backport_agent.pyClear matching backport failures after success +7/-0

Clear matching backport failures after success

• Successful backports remove prior error-list entries for the same issue, backport queue, and dist-git branch.

ymir/agents/backport_agent.py

rebase_agent.pyClear matching rebase failures after success +7/-0

Clear matching rebase failures after success

• Successful rebases remove prior error-list entries scoped to their issue, rebase queue, and dist-git branch.

ymir/agents/rebase_agent.py

rebuild_agent.pyClear matching rebuild failures after success +7/-0

Clear matching rebuild failures after success

• Successful rebuilds remove prior error-list entries scoped to their issue, rebuild queue, and dist-git branch.

ymir/agents/rebuild_agent.py

reproducer_agent.pyProcess requeued reproducers and clear successful failures +9/-0

Process requeued reproducers and clear successful failures

• The duplicate check permits tasks marked as requeued from the error list despite terminal labels. Successful reproducer results clear matching failures for the issue, workflow, and target branch.

ymir/agents/reproducer_agent.py

triage_agent.pyProcess requeued triage tasks and clear resolved failures +5/-0

Process requeued triage tasks and clear resolved failures

• The duplicate check permits marked requeues despite existing error labels. Non-error triage resolutions clear matching triage failures.

ymir/agents/triage_agent.py

error_list.pyAdd scoped cleanup for resolved error-list entries +46/-0

Add scoped cleanup for resolved error-list entries

• A shared helper matches failures by issue, normalized ordinary or priority queue, and target branch, then removes their original Redis list values. It leaves unmatchable entries intact and logs cleanup failures without failing completed workflows.

ymir/common/error_list.py

Tests (2) +87 / -0
test_requeue_error.pyTest normal and priority requeue preparation +29/-0

Test normal and priority requeue preparation

• New unit tests verify destination queues, reset attempts, the user-triggered setting, and the error-list requeue marker.

openshift/scripts/tests/unit/test_requeue_error.py

test_error_list.pyTest scoped cleanup and cleanup-failure handling +58/-0

Test scoped cleanup and cleanup-failure handling

• New tests verify that matching ordinary and priority entries are removed while other issues, workflows, branches, and legacy entries remain. They also verify that a Redis failure does not fail the completed workflow.

ymir/common/tests/unit/test_error_list.py

@qodo-for-packit

qodo-for-packit Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

Code Review by Qodo

🐞 Bugs (1) 📘 Rule violations (0) 📜 Skill insights (0)

Grey Divider


Action required

1. Failed reproducer updates erase prior errors ✓ Resolved
Description
_clear_finalized_reproducer_errors treats test_already_exists as resolved even when an attempted
update to that test has failed. Adapted-existing results set that flag, and merge-request failures
change success to false without clearing it, so a later cleanup removes matching records from
error_list despite the failed update.
Code

ymir/agents/reproducer_agent.py[R275-279]

+    if _determine_result_label(result) not in {
+        JiraLabels.REPRODUCER_CREATED,
+        JiraLabels.REPRODUCER_ALREADY_EXISTS,
+        JiraLabels.REPRODUCER_NOT_REPRODUCIBLE,
+    }:
Relevance

●●● Strong

Accepted patterns favor correcting failure-state cleanup; this directly contradicts the PR’s stated
requirement to keep failed reproducer errors.

PR-#843
PR-#611

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
The adapted-existing prompt sets success, test_already_exists, and adapted_existing to true.
Merge-request failure paths set only success to false; _determine_result_label then returns the
already-exists label, which the new cleanup accepts before clear_resolved_errors removes matching
Redis entries.

ymir/agents/prompts/reproducer/prompt.j2[159-177]
ymir/agents/reproducer_agent.py[962-979]
ymir/agents/reproducer_agent.py[235-243]
ymir/agents/reproducer_agent.py[269-287]
ymir/common/error_list.py[31-50]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
An adapted-existing reproducer can fail while updating its merge request yet retain `test_already_exists=True`; the new terminal-label check then deletes its prior error records.
## Fix Focus Areas
- ymir/agents/reproducer_agent.py[269-287]
- ymir/agents/tests/unit/test_reproducer_agent.py[117-162]
## Recommended Fix
Do not clear reproducer errors when `adapted_existing` is true and `success` is false, even if the result label is `REPRODUCER_ALREADY_EXISTS`. Add a regression test for the post-merge-request-failure flag combination.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


2. Dry runs delete active failure records ✓ Resolved
Description
The new clear_resolved_errors() calls are unconditional even though each queue-mode agent derives
a dry_run flag from DRY_RUN. A dry-run still consumes tasks from the live Redis queues, so a
successful simulated triage, rebase, backport, rebuild, or reproducer run removes matching
production retry records from error_list.
Code

ymir/agents/triage_agent.py[R1963-1964]

+                if output.resolution != Resolution.ERROR:
+                    await clear_resolved_errors(redis, input.issue, RedisQueues.TRIAGE_QUEUE.value)
Relevance

●●● Strong

Unconditional LREM mutates shared production state during dry runs, a clear correctness and safety
defect.

PR-#843

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
clear_resolved_errors removes Redis list entries with LREM, while the affected agents parse but
do not apply dry_run at the new call sites. The generic queue loop always pops work from Redis, so
dry-run execution can reach this mutation against existing shared state.

ymir/common/error_list.py[12-41]
ymir/common/base_utils.py[134-216]
ymir/agents/triage_agent.py[1361-1361]
ymir/agents/triage_agent.py[1963-1964]
ymir/agents/rebase_agent.py[156-156]
ymir/agents/rebase_agent.py[924-929]
ymir/agents/backport_agent.py[1966-1966]
ymir/agents/backport_agent.py[2195-2200]
ymir/agents/rebuild_agent.py[68-68]
ymir/agents/rebuild_agent.py[673-678]
ymir/agents/reproducer_agent.py[1167-1167]
ymir/agents/reproducer_agent.py[1444-1450]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

Issue description
The resolved-error cleanup mutates Redis even when an agent is running with `DRY_RUN=true`. Queue-mode dry runs consume live tasks, so they must not remove active production error-list records.

Fix Focus Areas
- ymir/agents/triage_agent.py[1963-1964]
- ymir/agents/rebase_agent.py[924-929]
- ymir/agents/backport_agent.py[2195-2200]
- ymir/agents/rebuild_agent.py[673-678]
- ymir/agents/reproducer_agent.py[1444-1450]

Recommended Fix
Guard every `clear_resolved_errors()` invocation with `if not dry_run` (preserving the reproducer's resolved-outcome condition as well), or add an explicit dry-run no-op to the helper and pass the flag from every caller. Add tests proving successful dry-run processing does not call `LREM`.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools



Remediation recommended

3. Success cleanup slows as failures pile up 🐞 Bug ➹ Performance ⭐ New
Description
clear_resolved_errors loads the whole error_list with LRANGE 0 -1 and runs json.loads on every
entry each time a triage, backport, rebase, rebuild or reproducer run succeeds. Nothing trims the
list, and each entry holds a full traceback plus the task metadata. The rebase and rebuild helpers
call it once per consolidated issue, so the same full list is read several times per success.
Code

ymir/common/error_list.py[R34-37]

+        entries = await fix_await(redis_conn.lrange(RedisQueues.ERROR_LIST.value, 0, -1))
+        for raw in entries:
+            try:
+                entry = json.loads(raw)
Relevance

●● Moderate

Concrete unbounded Redis scan, but remedy requires architectural redesign beyond this
cleanup-focused PR.

PR-#675
PR-#843

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
The cleanup reads the whole list on each success. Entries are only ever added: there are lpush calls
to ERROR_LIST in each agent and in tasks.py, and a search for ltrim finds nothing. Each entry stores
a traceback and the full task metadata, and the consolidated helpers call the full read once per
issue key.

ymir/common/error_list.py[34-50]
ymir/agents/rebase_agent.py[112-119]
ymir/agents/rebuild_agent.py[69-76]
ymir/agents/rebase_agent.py[863-866]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
`clear_resolved_errors` reads and parses the entire, never-trimmed `error_list` on every successful workflow. The rebase and rebuild helpers call it once per consolidated issue, so they repeat that full read.

## Fix Focus Areas
- ymir/common/error_list.py[12-55]
- ymir/agents/rebase_agent.py[103-119]
- ymir/agents/rebuild_agent.py[60-76]

## Recommended Fix
Change `clear_resolved_errors` to accept a collection of issue keys, or add a variant that does. Call LRANGE once, match each entry against the set of keys (plus queue and branch), and LREM the matching raw values. Update `_clear_rebase_resolved_errors` and `_clear_rebuild_resolved_errors` to make a single call with all deduplicated keys. Optionally cap the list's size, or keep a per-issue index so matching entries can be found without a full scan.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools

Dismiss ↗ | View ↗


4. Sibling rebase failures stay listed ✓ Resolved
Description
rebase_agent.retry clears resolved errors only for rebase_data.jira_issue, although a successful
rebase also resolves its consolidated siblings. When a sibling has an earlier rebase error for the
same branch, that entry remains in error_list after the sibling receives a success label.
Code

ymir/agents/rebase_agent.py[926]

+                        rebase_data.jira_issue,
Relevance

●●● Strong

Accepted precedent favors fixing consolidated rebase sibling handling; cleanup currently targets
only the primary issue.

PR-#726
PR-#785

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
The success path applies rebase success labels to consolidated issues, but the new cleanup passes
only the primary key. The helper requires an exact match on the error entry's Jira issue, so it
cannot remove a sibling's earlier entry.

ymir/agents/rebase_agent.py[899-930]
ymir/common/models.py[325-327]
ymir/common/error_list.py[42-50]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
A successful consolidated rebase clears prior errors for its primary issue but not its resolved sibling issues.

## Fix Focus Areas
- ymir/agents/rebase_agent.py[924-930]
- ymir/common/error_list.py[42-50]

## Recommended Fix
After successful processing, call `clear_resolved_errors` for the primary issue and each distinct consolidated sibling, retaining the workflow queue, target branch, and dry-run arguments.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


5. Sibling rebuild failures stay listed ✓ Resolved
Description
rebuild_agent.retry clears resolved errors only for rebuild_data.jira_issue, while its success
path marks every issue in rebuild_data.all_jira_issues as rebuilt. If a consolidated sibling has
an earlier rebuild error for that branch, its entry survives the successful rebuild.
Code

ymir/agents/rebuild_agent.py[675]

+                        rebuild_data.jira_issue,
Relevance

●●● Strong

Accepted precedent favors correcting consolidated rebuild behavior; successful sibling processing
should clear sibling failures too.

PR-#611
PR-#785

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
The rebuild success path labels all consolidated issues as rebuilt, but the new cleanup supplies
only the primary issue. Because cleanup matches error.jira_issue exactly, it leaves any sibling's
matching rebuild failure record untouched.

ymir/agents/rebuild_agent.py[647-679]
ymir/common/models.py[425-427]
ymir/common/error_list.py[42-50]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
A successful consolidated rebuild leaves prior error-list entries belonging to resolved sibling issues.

## Fix Focus Areas
- ymir/agents/rebuild_agent.py[647-679]
- ymir/common/error_list.py[42-50]

## Recommended Fix
Run `clear_resolved_errors` for each distinct issue in `rebuild_data.all_jira_issues` after successful completion, using the existing queue, branch, and dry-run arguments.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


View medium (2)
6. Requeued triage can skip reproducer reruns ✓ Resolved
Description
prepare_requeue_task marks a normal triage requeue, but _enqueue_reproducer creates a new task
without carrying that marker or setting user_triggered. If the issue already has a terminal
reproducer label, triage proceeds while the reproducer's duplicate guard discards the downstream
task.
Code

openshift/scripts/requeue_error.py[R172-173]

+    task["user_triggered"] = user_triggered
+    task["requeued_from_error_list"] = True
Relevance

●●● Strong

Directly breaks the PR’s requeue intent; downstream task construction must preserve the bypass
marker.

PR-#589

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
The requeue helper sets the marker while leaving user_triggered false for a normal run. Triage
uses the marker to bypass its duplicate guard, but its downstream task construction omits it; the
reproducer guard then rejects that task when a terminal reproducer label exists.

openshift/scripts/requeue_error.py[169-178]
ymir/agents/triage_agent.py[1525-1539]
ymir/agents/triage_agent.py[229-248]
ymir/agents/triage_agent.py[1945-1962]
ymir/agents/reproducer_agent.py[1233-1245]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
A normal error-list requeue can run triage but have its downstream reproducer skipped when a terminal reproducer label remains.
## Fix Focus Areas
- openshift/scripts/requeue_error.py[169-178]
- ymir/agents/triage_agent.py[229-248]
- ymir/agents/triage_agent.py[1945-1962]
## Recommended Fix
Pass the requeued-from-error-list state into `_enqueue_reproducer` and set it on the new reproducer `Task`, without changing the normal queue's priority.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


7. Resolved reproducer errors stay listed ✓ Resolved
Description
retry calls clear_resolved_errors() only when output.success is true, even though
not-reproducible and already-exists results can receive terminal labels without it. When either
outcome follows a recorded failure for the same issue and branch, the matching entry remains in
error_list after the Jira label is written.
Code

ymir/agents/reproducer_agent.py[R1444-1445]

+                    if output.success:
+                        await clear_resolved_errors(
Relevance

●●● Strong

Terminal unsuccessful outcomes still resolve reproducer work, so cleanup should not depend solely on
success.

PR-#843

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
The result-label logic maps unsuccessful not-reproducible and already-exists outputs to terminal
labels, and Jira finalization does not require success. The cleanup guard nevertheless requires
output.success, so it skips those outcomes; retry-exhaustion errors carry the issue, queue, and
task metadata that the skipped helper would otherwise use to match the earlier entry.

ymir/agents/reproducer_agent.py[237-263]
ymir/agents/reproducer_agent.py[1101-1108]
ymir/agents/reproducer_agent.py[1275-1300]
ymir/agents/reproducer_agent.py[1434-1450]
ymir/common/error_list.py[22-41]
ymir/agents/reproducer_agent.py[234-256]
ymir/agents/reproducer_agent.py[260-267]
ymir/agents/reproducer_agent.py[1058-1117]
ymir/agents/reproducer_agent.py[1284-1304]
ymir/agents/reproducer_agent.py[1419-1450]
ymir/common/error_list.py[15-41]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
Finalized not-reproducible and already-exists reproducer outcomes can leave earlier matching failures in the active error list because cleanup depends on `output.success`.

## Fix Focus Areas
- ymir/agents/reproducer_agent.py[237-267]
- ymir/agents/reproducer_agent.py[1434-1450]

## Recommended Fix
Base matching error-list cleanup on the terminal result label rather than `output.success` alone. Call `clear_resolved_errors()` for `REPRODUCER_CREATED`, `REPRODUCER_ALREADY_EXISTS`, and `REPRODUCER_NOT_REPRODUCIBLE`, while retaining error records for `REPRODUCER_FAILED` and deferred or retryable results. Add a regression test for a false-success result with `not_reproducible_reason` following a recorded failure.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


Grey Divider

Context sources
✅ Compliance rules (platform): 8 rules
Review mode: 🧠 Deep: This push adds behavior across several independent workflow agents plus shared Redis error-list cleanup semantics, creating multiple plausible, easy-to-miss correctness paths that benefit from redundant review.

Grey Divider

Tip of the day
💡 Did you know, you can route each severity your way: inline, summary, both, or drop

More tips ↗ | Customize Qodo ↗ | Qodo docs ↗

Grey Divider

Previous reviews

Review updated until commit 40d8e94

Results up to commit 676a16a 🧠 Deep


🐞 Bugs (0) 📘 Rule violations (0) 📎 Requirement gaps (0) 🎨 UX issues (0) 🔗 Cross-repo conflicts (0) 📜 Skill insights (0)


Action required
1. Dry runs delete active failure records ✓ Resolved
Description
The new clear_resolved_errors() calls are unconditional even though each queue-mode agent derives
a dry_run flag from DRY_RUN. A dry-run still consumes tasks from the live Redis queues, so a
successful simulated triage, rebase, backport, rebuild, or reproducer run removes matching
production retry records from error_list.
Code

ymir/agents/triage_agent.py[R1963-1964]

+                if output.resolution != Resolution.ERROR:
+                    await clear_resolved_errors(redis, input.issue, RedisQueues.TRIAGE_QUEUE.value)
Relevance

●●● Strong

Unconditional LREM mutates shared production state during dry runs, a clear correctness and safety
defect.

PR-#843

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
clear_resolved_errors removes Redis list entries with LREM, while the affected agents parse but
do not apply dry_run at the new call sites. The generic queue loop always pops work from Redis, so
dry-run execution can reach this mutation against existing shared state.

ymir/common/error_list.py[12-41]
ymir/common/base_utils.py[134-216]
ymir/agents/triage_agent.py[1361-1361]
ymir/agents/triage_agent.py[1963-1964]
ymir/agents/rebase_agent.py[156-156]
ymir/agents/rebase_agent.py[924-929]
ymir/agents/backport_agent.py[1966-1966]
ymir/agents/backport_agent.py[2195-2200]
ymir/agents/rebuild_agent.py[68-68]
ymir/agents/rebuild_agent.py[673-678]
ymir/agents/reproducer_agent.py[1167-1167]
ymir/agents/reproducer_agent.py[1444-1450]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

Issue description
The resolved-error cleanup mutates Redis even when an agent is running with `DRY_RUN=true`. Queue-mode dry runs consume live tasks, so they must not remove active production error-list records.

Fix Focus Areas
- ymir/agents/triage_agent.py[1963-1964]
- ymir/agents/rebase_agent.py[924-929]
- ymir/agents/backport_agent.py[2195-2200]
- ymir/agents/rebuild_agent.py[673-678]
- ymir/agents/reproducer_agent.py[1444-1450]

Recommended Fix
Guard every `clear_resolved_errors()` invocation with `if not dry_run` (preserving the reproducer's resolved-outcome condition as well), or add an explicit dry-run no-op to the helper and pass the flag from every caller. Add tests proving successful dry-run processing does not call `LREM`.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools



Remediation recommended
2. Resolved reproducer errors stay listed ✓ Resolved
Description
retry calls clear_resolved_errors() only when output.success is true, even though
not-reproducible and already-exists results can receive terminal labels without it. When either
outcome follows a recorded failure for the same issue and branch, the matching entry remains in
error_list after the Jira label is written.
Code

ymir/agents/reproducer_agent.py[R1444-1445]

+                    if output.success:
+                        await clear_resolved_errors(
Relevance

●●● Strong

Terminal unsuccessful outcomes still resolve reproducer work, so cleanup should not depend solely on
success.

PR-#843

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
The result-label logic maps unsuccessful not-reproducible and already-exists outputs to terminal
labels, and Jira finalization does not require success. The cleanup guard nevertheless requires
output.success, so it skips those outcomes; retry-exhaustion errors carry the issue, queue, and
task metadata that the skipped helper would otherwise use to match the earlier entry.

ymir/agents/reproducer_agent.py[237-263]
ymir/agents/reproducer_agent.py[1101-1108]
ymir/agents/reproducer_agent.py[1275-1300]
ymir/agents/reproducer_agent.py[1434-1450]
ymir/common/error_list.py[22-41]
ymir/agents/reproducer_agent.py[234-256]
ymir/agents/reproducer_agent.py[260-267]
ymir/agents/reproducer_agent.py[1058-1117]
ymir/agents/reproducer_agent.py[1284-1304]
ymir/agents/reproducer_agent.py[1419-1450]
ymir/common/error_list.py[15-41]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
Finalized not-reproducible and already-exists reproducer outcomes can leave earlier matching failures in the active error list because cleanup depends on `output.success`.

## Fix Focus Areas
- ymir/agents/reproducer_agent.py[237-267]
- ymir/agents/reproducer_agent.py[1434-1450]

## Recommended Fix
Base matching error-list cleanup on the terminal result label rather than `output.success` alone. Call `clear_resolved_errors()` for `REPRODUCER_CREATED`, `REPRODUCER_ALREADY_EXISTS`, and `REPRODUCER_NOT_REPRODUCIBLE`, while retaining error records for `REPRODUCER_FAILED` and deferred or retryable results. Add a regression test for a false-success result with `not_reproducible_reason` following a recorded failure.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


3. Requeued triage can skip reproducer reruns ✓ Resolved
Description
prepare_requeue_task marks a normal triage requeue, but _enqueue_reproducer creates a new task
without carrying that marker or setting user_triggered. If the issue already has a terminal
reproducer label, triage proceeds while the reproducer's duplicate guard discards the downstream
task.
Code

openshift/scripts/requeue_error.py[R172-173]

+    task["user_triggered"] = user_triggered
+    task["requeued_from_error_list"] = True
Relevance

●●● Strong

Directly breaks the PR’s requeue intent; downstream task construction must preserve the bypass
marker.

PR-#589

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
The requeue helper sets the marker while leaving user_triggered false for a normal run. Triage
uses the marker to bypass its duplicate guard, but its downstream task construction omits it; the
reproducer guard then rejects that task when a terminal reproducer label exists.

openshift/scripts/requeue_error.py[169-178]
ymir/agents/triage_agent.py[1525-1539]
ymir/agents/triage_agent.py[229-248]
ymir/agents/triage_agent.py[1945-1962]
ymir/agents/reproducer_agent.py[1233-1245]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
A normal error-list requeue can run triage but have its downstream reproducer skipped when a terminal reproducer label remains.
## Fix Focus Areas
- openshift/scripts/requeue_error.py[169-178]
- ymir/agents/triage_agent.py[229-248]
- ymir/agents/triage_agent.py[1945-1962]
## Recommended Fix
Pass the requeued-from-error-list state into `_enqueue_reproducer` and set it on the new reproducer `Task`, without changing the normal queue's priority.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


Results up to commit 77b3d45 ⚖️ Balanced


🐞 Bugs (0) 📘 Rule violations (0) 📎 Requirement gaps (0) 🎨 UX issues (0) 🔗 Cross-repo conflicts (0) 📜 Skill insights (0)


Remediation recommended
1. Sibling rebase failures stay listed ✓ Resolved
Description
rebase_agent.retry clears resolved errors only for rebase_data.jira_issue, although a successful
rebase also resolves its consolidated siblings. When a sibling has an earlier rebase error for the
same branch, that entry remains in error_list after the sibling receives a success label.
Code

ymir/agents/rebase_agent.py[926]

+                        rebase_data.jira_issue,
Relevance

●●● Strong

Accepted precedent favors fixing consolidated rebase sibling handling; cleanup currently targets
only the primary issue.

PR-#726
PR-#785

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
The success path applies rebase success labels to consolidated issues, but the new cleanup passes
only the primary key. The helper requires an exact match on the error entry's Jira issue, so it
cannot remove a sibling's earlier entry.

ymir/agents/rebase_agent.py[899-930]
ymir/common/models.py[325-327]
ymir/common/error_list.py[42-50]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
A successful consolidated rebase clears prior errors for its primary issue but not its resolved sibling issues.

## Fix Focus Areas
- ymir/agents/rebase_agent.py[924-930]
- ymir/common/error_list.py[42-50]

## Recommended Fix
After successful processing, call `clear_resolved_errors` for the primary issue and each distinct consolidated sibling, retaining the workflow queue, target branch, and dry-run arguments.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


2. Sibling rebuild failures stay listed ✓ Resolved
Description
rebuild_agent.retry clears resolved errors only for rebuild_data.jira_issue, while its success
path marks every issue in rebuild_data.all_jira_issues as rebuilt. If a consolidated sibling has
an earlier rebuild error for that branch, its entry survives the successful rebuild.
Code

ymir/agents/rebuild_agent.py[675]

+                        rebuild_data.jira_issue,
Relevance

●●● Strong

Accepted precedent favors correcting consolidated rebuild behavior; successful sibling processing
should clear sibling failures too.

PR-#611
PR-#785

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
The rebuild success path labels all consolidated issues as rebuilt, but the new cleanup supplies
only the primary issue. Because cleanup matches error.jira_issue exactly, it leaves any sibling's
matching rebuild failure record untouched.

ymir/agents/rebuild_agent.py[647-679]
ymir/common/models.py[425-427]
ymir/common/error_list.py[42-50]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
A successful consolidated rebuild leaves prior error-list entries belonging to resolved sibling issues.

## Fix Focus Areas
- ymir/agents/rebuild_agent.py[647-679]
- ymir/common/error_list.py[42-50]

## Recommended Fix
Run `clear_resolved_errors` for each distinct issue in `rebuild_data.all_jira_issues` after successful completion, using the existing queue, branch, and dry-run arguments.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


Results up to commit 21ce79d ⚖️ Balanced


🐞 Bugs (0) 📘 Rule violations (0) 📎 Requirement gaps (0) 🎨 UX issues (0) 🔗 Cross-repo conflicts (0) 📜 Skill insights (0)


Action required
1. Failed reproducer updates erase prior errors ✓ Resolved
Description
_clear_finalized_reproducer_errors treats test_already_exists as resolved even when an attempted
update to that test has failed. Adapted-existing results set that flag, and merge-request failures
change success to false without clearing it, so a later cleanup removes matching records from
error_list despite the failed update.
Code

ymir/agents/reproducer_agent.py[R275-279]

+    if _determine_result_label(result) not in {
+        JiraLabels.REPRODUCER_CREATED,
+        JiraLabels.REPRODUCER_ALREADY_EXISTS,
+        JiraLabels.REPRODUCER_NOT_REPRODUCIBLE,
+    }:
Relevance

●●● Strong

Accepted patterns favor correcting failure-state cleanup; this directly contradicts the PR’s stated
requirement to keep failed reproducer errors.

PR-#843
PR-#611

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
The adapted-existing prompt sets success, test_already_exists, and adapted_existing to true.
Merge-request failure paths set only success to false; _determine_result_label then returns the
already-exists label, which the new cleanup accepts before clear_resolved_errors removes matching
Redis entries.

ymir/agents/prompts/reproducer/prompt.j2[159-177]
ymir/agents/reproducer_agent.py[962-979]
ymir/agents/reproducer_agent.py[235-243]
ymir/agents/reproducer_agent.py[269-287]
ymir/common/error_list.py[31-50]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
An adapted-existing reproducer can fail while updating its merge request yet retain `test_already_exists=True`; the new terminal-label check then deletes its prior error records.
## Fix Focus Areas
- ymir/agents/reproducer_agent.py[269-287]
- ymir/agents/tests/unit/test_reproducer_agent.py[117-162]
## Recommended Fix
Do not clear reproducer errors when `adapted_existing` is true and `success` is false, even if the result label is `REPRODUCER_ALREADY_EXISTS`. Add a regression test for the post-merge-request-failure flag combination.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


Grey Divider

Qodo Logo

Comment thread openshift/scripts/requeue_error.py
Comment thread ymir/agents/reproducer_agent.py Outdated
Comment thread ymir/agents/triage_agent.py Outdated
@majamassarini
majamassarini force-pushed the fix/error-list-requeue-and-cleanup branch from 676a16a to 8468c46 Compare September 28, 2026 12:57
@majamassarini majamassarini changed the title Fix error requeue routing and clear resolved failures Prioritize error requeues with less maintainer noise and clear resolved failures Sep 28, 2026
@majamassarini
majamassarini force-pushed the fix/error-list-requeue-and-cleanup branch 2 times, most recently from 27889ad to 77b3d45 Compare September 28, 2026 13:13
@majamassarini

Copy link
Copy Markdown
Member Author

/agentic_review

Comment thread ymir/agents/rebase_agent.py Outdated
Comment thread ymir/agents/rebuild_agent.py Outdated
@qodo-for-packit

Copy link
Copy Markdown

Code review by qodo was updated up to the latest commit 77b3d45

@majamassarini
majamassarini force-pushed the fix/error-list-requeue-and-cleanup branch 2 times, most recently from 1066a53 to 21ce79d Compare September 28, 2026 13:39
@majamassarini

Copy link
Copy Markdown
Member Author

/agentic_review

Comment thread ymir/agents/reproducer_agent.py
@qodo-for-packit

Copy link
Copy Markdown

Code review by qodo was updated up to the latest commit 21ce79d

@majamassarini
majamassarini force-pushed the fix/error-list-requeue-and-cleanup branch from 21ce79d to 19608e8 Compare September 29, 2026 08:10
@majamassarini

Copy link
Copy Markdown
Member Author

/agentic_review

Comment thread ymir/common/error_list.py
@qodo-for-packit

Copy link
Copy Markdown

Code review by qodo was updated up to the latest commit 19608e8

@majamassarini
majamassarini force-pushed the fix/error-list-requeue-and-cleanup branch from 19608e8 to 729cf73 Compare September 29, 2026 08:53
@majamassarini
majamassarini force-pushed the fix/error-list-requeue-and-cleanup branch from 729cf73 to 40d8e94 Compare September 30, 2026 09:36
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant