Skip to content

fix(workflow): keep finished parallel tool results when a sibling raises - #7429

Open
thuongvu wants to merge 2 commits into
google:mainfrom
thuongvu:fix/keep-completed-parallel-results-on-tool-error
Open

thuongvu wants to merge 2 commits into
google:mainfrom
thuongvu:fix/keep-completed-parallel-results-on-tool-error

Conversation

@thuongvu

@thuongvu thuongvu commented Oct 6, 2026

Copy link
Copy Markdown

Link to Issue or Description of Change

1. Link to an existing issue (if applicable):

Problem:

When one of several parallel tool calls raises, the batch is canceled and the error re-raised, but a sibling that had
already finished loses its response with the batch. It's then stripped from the next request, so the model doesn't
know it ran and may repeat it.

Solution:

  • When the batch fails, the batch executor keeps the finished calls' responses (merged as a complete batch would be) on
    the invocation's shared _AbortState.
  • Once the node has stopped on the error, NodeRunner persists them through the node's normal event path, before it
    emits the error event or retries.
  • The original error then propagates unchanged.

Nothing is yielded into the agent while the error is in flight, so nothing can act on the kept results. Only plain
results are kept: responses with control actions (transfer, escalate, finish_task) or confirmation/credential
placeholders are left out, so history only records what actually ran. Live mode and aborted invocations are
unchanged (abort sealing answers those calls, as today).

Behavior changes to be aware of:

  • The kept event is the normal merged response event, so its state_delta / artifact_delta are now persisted too.
  • If a failed invocation is resumed, the model now sees the finished call's result instead of the call being re-run.

Testing Plan

Unit Tests:

  • I have added or updated unit tests for my change.
  • All unit tests pass locally.

New tests in test_runners.py and test_batch_executor.py. Without the fix 13 of the 15 fail (the other 2 are
regression guards).

$ pytest tests/unittests/test_runners.py tests/unittests/flows/llm_flows/tools/test_batch_executor.py \
    -k "tool_error or keep_completed or plain_result"
15 passed, 194 deselected, 4 warnings in 2.22s

$ pytest tests/unittests -n 4
17832 passed, 87 skipped, 25 xfailed, 2 xpassed, 2320 warnings, 28 subtests passed in 101.39s

(The full run deselects cli/utils/test_cli_create.py::test_handle_login_with_google_option_3, which opens an
interactive Google login.)

Manual End-to-End (E2E) Tests:

The repro script from #7428:

$ python repro.py            # google-adk 2.11.0
raised: pager service returned 500
tickets created: ['disk full']
responses in session: []

$ python repro.py            # this branch
raised: pager service returned 500
tickets created: ['disk full']
responses in session: [{'ticket_id': 'T-1'}]

With a real model (gpt-5-mini via LiteLlm) asked to finish the task, the finished write was repeated in 10/10 runs
before this change and 0/10 after.

Checklist

  • I have read the CONTRIBUTING.md document.
  • I have performed a self-review of my own code.
  • I have commented my code, particularly in hard-to-understand areas.
  • I have added tests that prove my fix is effective or that my feature works.
  • New and existing unit tests pass locally with my changes.
  • I have manually tested my changes end-to-end.
  • Any dependent changes have been merged and published in downstream modules. (N/A)

Additional context

When the model emits parallel function calls and one FunctionTool raises,
the batch is cancelled and the error re-raised, but the response of a
sibling that had already finished was thrown away with it. The call was
then stripped from the next request, so the model could repeat it.

The batch executor now keeps the finished calls' real responses, and once
the node has stopped on the error, the node runner persists them before
the error event or a retry; the original error propagates unchanged.
Nothing is synthesized, and responses that carry control actions or
confirmation/credential placeholders are not kept.
@thuongvu
thuongvu force-pushed the fix/keep-completed-parallel-results-on-tool-error branch from fde7957 to 3e309d5 Compare October 9, 2026 00:06
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Parallel function calls: when one tool raises, finished siblings lose their responses and may be re-run

1 participant