Skip to content

fix(litellm): keep a streamed thought signature sent without the call's name - #7467

Open
vetler wants to merge 2 commits into
google:mainfrom
vetler:fix/litellm-signature-only-tool-call-delta
Open

vetler wants to merge 2 commits into
google:mainfrom
vetler:fix/litellm-signature-only-tool-call-delta

Conversation

@vetler

@vetler vetler commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

Please ensure you have read the contribution guide before creating a pull request.

Link to Issue or Description of Change

1. Link to an existing issue (if applicable):

This is the follow-up to #7441 for the case @surajksharma07 raised on #7438: a streamed thought signature that arrives on a tool-call delta with no name or arguments. #7441 landed without it (3abea3e); #7485 tracks it.

Problem:

_model_response_to_chunk drops a tool-call delta that has neither a name nor arguments as empty. A provider that sends the thought signature on such a delta, apart from the call's name and arguments, therefore still loses it, and Gemini 3 rejects the next request in the turn ("Function call is missing a thought_signature in functionCall parts").

Solution:

  • A delta that carries a signature is no longer treated as empty, even without a name or arguments.
  • In the streaming loop, such a delta attaches its signature to the call at its own index. The fallback index can already point past that call, so it is not used for the lookup.
  • If that call has not started yet, the signature is held and set when the call's first named chunk arrives.
  • The delta never opens a call of its own, which would reach the model as a call with no name, and it yields no partial response.

In live runs with gemini-3-flash-preview, neither Vertex AI's OpenAI-compatible endpoint nor LiteLLM's vertex_ai provider sent this shape: both put the signature on the delta that names the call (raw deltas below). So the unit tests are what exercise the new path; the live runs show the existing paths are unchanged.

Testing Plan

Unit Tests:

  • I have added or updated unit tests for my change.
  • All unit tests pass locally.

New tests in tests/unittests/models/test_litellm.py, all failing on main and passing with this change:

  • test_model_response_to_chunk_keeps_a_signature_only_delta: a delta with only a signature is no longer dropped as empty.
  • test_streaming_signature_after_its_call_signs_that_call: a signature-only delta that arrives after its call, with a second call in between, signs the first call and not the second.
  • test_streaming_signature_before_its_call_signs_that_call: a signature-only delta ahead of its call opens no call of its own and signs the call that follows.
$ pytest tests/unittests/models/ tests/unittests/flows -q
2597 passed

$ tox -p auto   # pytest tests/unittests on Python 3.11-3.14
py311: 18307 passed, 4 failed
py312: 18298 passed, 4 failed
py313: 18298 passed, 4 failed
py314: 18298 passed, 4 failed

The four failures are the TestParseToolCallArguments logging tests, which depend on test order and fail the same way on main (fix in #7461).

pre-commit run --files passes on both changed files, and mypy reports no errors in lite_llm.py.

Manual End-to-End (E2E) Tests:

The same live check as on #7441: gemini-3-flash-preview through a Runner with RunConfig(streaming_mode=StreamingMode.SSE), asked for the weather and the local time in two cities so the model makes four parallel calls, with every raw tool-call delta logged. The script is in #7441's description (VERTEX_PROJECT=<project> ROUTE=openai|vertex_ai LOG_DELTAS=1 python e2e_parallel.py).

OpenAI-compatible endpoint:

[request 1 delta 1] index=0 id=call_2175506 name='get_weather' args_len=15 extra_content.google.thought_signature=yes provider_specific_fields.thought_signature=no finish=None
[request 1 delta 2] index=1 id=call_2175509 name='get_time' args_len=15 extra_content.google.thought_signature=no provider_specific_fields.thought_signature=no finish=None
[request 1 delta 3] index=2 id=call_2175517 name='get_weather' args_len=17 extra_content.google.thought_signature=no provider_specific_fields.thought_signature=no finish=None
[request 1 delta 4] index=3 id=call_2175518 name='get_time' args_len=17 extra_content.google.thought_signature=no provider_specific_fields.thought_signature=no finish=None
[assistant] function_call get_weather({'city': 'Oslo'}) thought_signature=<426 bytes>
[assistant] function_call get_time({'city': 'Oslo'}) thought_signature=None
[assistant] function_call get_weather({'city': 'Bergen'}) thought_signature=None
[assistant] function_call get_time({'city': 'Bergen'}) thought_signature=None
[assistant] function_response get_weather
[assistant] function_response get_time
[assistant] function_response get_weather
[assistant] function_response get_time
[assistant] text: In Oslo, it is currently 14:05 and the weather is rainy with a temperature of 9°C. [rest of the answer, about Bergen, trimmed]

LiteLLM vertex_ai provider:

[request 1 delta 1] index=0 id=<embedded sig> name='get_weather' args_len=16 extra_content.google.thought_signature=no provider_specific_fields.thought_signature=yes finish=None
[request 1 delta 2] index=1 id=call_3868561 name='get_time' args_len=16 extra_content.google.thought_signature=no provider_specific_fields.thought_signature=no finish=None
[request 1 delta 3] index=2 id=call_3868564 name='get_weather' args_len=18 extra_content.google.thought_signature=no provider_specific_fields.thought_signature=no finish=None
[request 1 delta 4] index=3 id=call_3868567 name='get_time' args_len=18 extra_content.google.thought_signature=no provider_specific_fields.thought_signature=no finish=None
[assistant] function_call get_weather({'city': 'Oslo'}) thought_signature=<658 bytes>
[assistant] function_call get_time({'city': 'Oslo'}) thought_signature=None
[assistant] function_call get_weather({'city': 'Bergen'}) thought_signature=None
[assistant] function_call get_time({'city': 'Bergen'}) thought_signature=None
[assistant] function_response get_weather
[assistant] function_response get_time
[assistant] function_response get_weather
[assistant] function_response get_time
[assistant] text: In Oslo, it is currently 14:05 and the weather is rainy with a temperature of 9°C. Similarly, in Bergen, it is 14:05 and the weather is also rainy with a temperature of 9°C.

Checklist

  • I have read the CONTRIBUTING.md document.
  • I have performed a self-review of my own code.
  • I have commented my code, particularly in hard-to-understand areas.
  • I have added tests that prove my fix is effective or that my feature works.
  • New and existing unit tests pass locally with my changes.
  • I have manually tested my changes end-to-end.
  • Any dependent changes have been merged and published in downstream modules.

🤖 Generated with Claude Code

…'s name

A tool-call delta carrying only a signature was dropped as empty, so a
provider that sends the signature apart from the name and arguments would
still lose it. Keep such deltas, attach the signature to the call at the
delta's own index, or hold it for the call that starts there next, and never
open a call for it, which would reach the model without a name.

In live runs with gemini-3-flash-preview, neither Vertex AI's
OpenAI-compatible endpoint nor LiteLLM's vertex_ai provider sent this
shape; both put the signature on the delta that names the call.

Follow-up to google#7441, for the case raised on google#7438.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

LiteLlm streaming drops a thought_signature sent on a tool-call delta without the call's name or arguments

2 participants