Skip to content

LiteLlm streaming drops Gemini thought_signature from tool calls; Gemini 3 then rejects the follow-up request聽#7438

Description

@vetler

馃敶 Required Information

Describe the Bug:
With LiteLlm and stream=True (for example /run_sse with streaming: true), the thought_signature that Gemini 3 attaches to a tool call is dropped while the streamed tool call is reassembled. The function-call part ADK stores has no signature, so when ADK sends the history back after the tool runs, Gemini 3 rejects the request:

Function call is missing a thought_signature in functionCall parts. This is required for tools to work correctly, and missing thought_signature may lead to degraded model performance.

Without streaming the signature is kept. The non-streaming conversion paths were fixed in #3627 and #4650; the streaming path still loses it.

Steps to Reproduce:

  1. pip install google-adk litellm (also reproduces on main at 77d4dcd).
  2. Run the script under Minimal Reproduction Code. It mocks the LiteLLM client with the shape the Vertex AI OpenAI-compatible endpoint returns, so no credentials are needed.
  3. Compare the two printed lines.

Against a live endpoint: LiteLlm(model="google/gemini-3-flash-preview", api_base="https://aiplatform.googleapis.com/v1/projects/<project>/locations/global/endpoints/openapi", custom_llm_provider="openai") with any tool, stream=True. The request after the tool call fails with the 400 above.

Expected Behavior:
The streamed tool call keeps its signature, as the non-streamed one does:

non-streamed thought_signature: b'opaque-signature'
streamed thought_signature:     b'opaque-signature'

Observed Behavior:

non-streamed thought_signature: b'opaque-signature'
streamed thought_signature:     None

Live, the follow-up request returns the 400 quoted above.

Environment Details:

  • ADK Library Version (pip show google-adk): 2.11.0, and main at 77d4dcd
  • Desktop OS: macOS
  • Python Version (python -V): 3.14.7

Model Information:

  • Are you using LiteLLM: Yes (litellm 1.104.0)
  • Which model is being used: Gemini 3 Flash models on the Vertex AI OpenAI-compatible endpoint (gemini-3.8-flash, gemini-3-flash-preview)

馃煛 Optional Information

Regression:
No; 2.5.0 drops it as well.

Additional Context:
LiteLLM keeps the signature on the streamed delta: tool_call.get("extra_content") returns {"google": {"thought_signature": "..."}}. It is lost in ADK:

Possible fix: add an optional thought_signature to FunctionChunk, set it from _extract_thought_signature_from_tool_call(tool_call) in _model_response_to_chunk, keep it per index in function_calls, and put it back as extra_content={"google": {"thought_signature": ...}} on the assembled ChatCompletionMessageToolCall. With parallel calls Gemini signs only the first, so the field must stay optional.

Why it can go unnoticed: Gemini 3 checks signatures only for function calls in the current turn. If a user-role message follows the function response, for example an instruction that ADK adds as user content when static_instruction is set, the call no longer counts as the current turn and the request succeeds, though without the signature.

Related: #7004 touches the same streamed tool-call aggregation.

Minimal Reproduction Code:

"""LiteLlm keeps a Gemini thought_signature on non-streamed tool calls but drops it when streaming.

No network or credentials: the LiteLLM client is mocked with the shape the
Vertex AI OpenAI-compatible endpoint returns for Gemini 3 tool calls.
"""

import asyncio
import base64
from unittest.mock import AsyncMock, MagicMock

import litellm
from google.adk.models.lite_llm import LiteLlm
from google.adk.models.llm_request import LlmRequest
from google.genai import types

SIGNATURE = base64.b64encode(b"opaque-signature").decode()
TOOL_CALL = {
    "index": 0,
    "id": "call_1",
    "type": "function",
    "function": {"name": "get_weather", "arguments": '{"city": "Oslo"}'},
    # Where Vertex puts the signature, in both streamed and non-streamed responses.
    "extra_content": {"google": {"thought_signature": SIGNATURE}},
}


def non_streamed_response():
    return litellm.ModelResponse(
        choices=[
            {
                "index": 0,
                "finish_reason": "tool_calls",
                "message": {"role": "assistant", "content": None, "tool_calls": [TOOL_CALL]},
            }
        ]
    )


async def streamed_response():
    yield litellm.ModelResponseStream(
        choices=[
            {
                "index": 0,
                "finish_reason": None,
                "delta": {"role": "assistant", "tool_calls": [TOOL_CALL]},
            }
        ]
    )
    yield litellm.ModelResponseStream(
        choices=[{"index": 0, "finish_reason": "tool_calls", "delta": {}}]
    )


async def function_call_signature(stream: bool):
    llm = LiteLlm(model="google/gemini-3-flash-preview", custom_llm_provider="openai")
    llm.llm_client = MagicMock()
    llm.llm_client.acompletion = AsyncMock(
        return_value=streamed_response() if stream else non_streamed_response()
    )
    request = LlmRequest(
        model="google/gemini-3-flash-preview",
        contents=[types.Content(role="user", parts=[types.Part(text="Weather in Oslo?")])],
        config=types.GenerateContentConfig(),
    )
    final = None
    async for response in llm.generate_content_async(request, stream=stream):
        if not response.partial:
            final = response
    part = next(p for p in final.content.parts if p.function_call)
    return part.thought_signature


async def main():
    print("non-streamed thought_signature:", await function_call_signature(stream=False))
    print("streamed thought_signature:    ", await function_call_signature(stream=True))


asyncio.run(main())

How often has this issue occurred?:

  • Always (100%) when streaming

Activity

  1. reddynitish commented on Oct 7, 2026

    @reddynitish

    I鈥檓 interested in working on this, if it is still available. My proposed scope is to preserve the optional thought signature through FunctionChunk and per-index streamed tool-call assembly, with mocked regression coverage for signed and unsigned calls and multiple tool-call indices. I checked the current streaming assembly and it does not carry the signature through.

    I noticed #7005 touches the same aggregation area, so I would coordinate with that work rather than duplicate or rewrite its fix. I found no public claim or PR for this specific signature-loss issue. If you or another contributor already have a fix underway, please let me know. Maintainers: would this scope be welcome?

    I plan to use Codex assistance. This is an interest/coordination request; I鈥檓 prioritizing one contribution at a time rather than promising simultaneous delivery.

  2. vetler commented on Oct 7, 2026

    @vetler
    ContributorAuthor

    @reddynitish I started working on a fix right before I saw your comment. You're welcome to review and contribute, or suggest another fix of course

  3. added a commit that references this issue on Oct 7, 2026
    42a5da3
  4. surajksharma07 commented on Oct 8, 2026

    @surajksharma07
    Collaborator

    @vetler Reproduced this on 2.11.0 and current main: the signature survives non-streaming but comes back None when streamed, for both extra_content and provider_specific_fields (an id-embedded signature does make it through).

    Tried carrying the signature per tool call through the stream and re-attaching it as extra_content on the rebuilt call, same idea as #7441 and it keeps the signature on the part and on the follow-up request including a parallel batch where only the first call is signed. Until that's merged turning streaming off is the quickest workaround.

    Could you check that this holds up against live Gemini 3 with a couple of parallel calls and also the case where the signature arrives on a delta with no name/args (those chunks are still skipped)? Please run the full unit suite on #7441 once more after that before it goes for review.

    #7450 targets the same fix so it'd be good to settle on one PR and #7005 touches the same loop so whichever lands second will need a rebase.

  5. vetler commented on Oct 8, 2026

    @vetler
    ContributorAuthor

    @surajksharma07 Thanks for reproducing it. I've pushed the follow-ups to #7441.

    Parallel calls against live Gemini 3. I asked gemini-3-flash-preview for the weather and the local time in two cities, which gives four parallel calls in one step, and logged every streamed tool-call delta:

    • OpenAI-compatible endpoint: the signature comes in extra_content.google on the first call only. main gets the 400 on the follow-up request. With fix(litellm): keep Gemini thought signatures on streamed tool calls聽#7441 the first call keeps its signature, the other three stay unsigned, and the turn completes.
    • LiteLLM's vertex_ai provider: the signature comes in provider_specific_fields and the call id. That already works on main and still does.

    Signature on a delta with no name or arguments. Neither route actually sends one, but since those deltas used to be dropped as empty, #7441 now keeps them. The signature goes to the call at the delta's own index, or waits until that call starts, and it never opens a call of its own. Tests cover both orders.

    Full unit suite on 3.11 to 3.14, with current main merged in: everything passes except failures main has as well. Those are the four TestParseToolCallArguments logging tests that main's CI fails on today (they break after test_agent_to_a2a.py, because to_a2a() sets the google_adk logger to INFO), plus one timing flake on 3.12 that passed on rerun. I pinned kubernetes to 36.0.3 for the GKE tests (#7443).

    #7450 and #7005. #7441 was opened first and only touches lite_llm.py and its tests, but I'm happy to go with whichever PR you prefer and bring over anything #7450 has that #7441 doesn't. If #7005 lands first, I'll update #7441 against main.

  6. surajksharma07 commented on Oct 8, 2026

    @surajksharma07
    Collaborator

    Thanks for the live parallel-call logs @vetler. #7441 landed as 3abea3e and re-ran the repro on main where the streamed signature now comes through. Not in a release yet (latest is v2.11.0) so until the next one either install from main or keep streaming off.

    One heads-up: the import picked up the earlier revision so c6882de (the signature-only delta handling and its three tests) didn't make it in. Main still drops those deltas and a signature that arrives on its own after the call still ends up None. Neither route sends that shape today so nobody is hitting it but if you want it in a small follow-up PR off main with just that commit should be quick to review.

    #7005 still applies cleanly on top of this so no rebase needed there. @sushant-me #7450 now duplicates what's merged so it can be closed. The a2a/tools/gke commits in it would be better off as separate PRs.

  7. added a commit that references this issue on Oct 8, 2026
    3abea3e
  8. added
    models[Component] This issue is related to model support
    on Oct 9, 2026
  9. vetler commented on Oct 9, 2026

    @vetler
    ContributorAuthor

    Thanks! Opened #7467 off main with just that commit and its tests.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

models[Component] This issue is related to model support

Type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions