Repository navigation
Live inference_call_count_v1 changes with text chunk boundaries #7351
Description
Activity
Reproduced this on 2.10.0 with your script and getting [1.0, 2.0, 3.0] @yang0228.
From what I can see every partial chunk carries model_version but the merged text from __build_full_text_response() doesn't and the eval conversion never looks at event.partial. A workaround worth trying: stamp model_version on the non-partial final text and skip partial events before convert_events_to_eval_invocations(). Locally that gave 1.0 for 1/2/3 chunks and still 2.0 for two real model turns in the same invocation.
Could you check whether that approach holds up on your side especially for the sub-agent, tool-call and usage cases you listed? Please test it thoroughly before raising the PR.
Will also ask the eval owners to confirm the counting semantics and it'd be good to line this up with #7323 so the two changes don't conflict.
Thanks @surajksharma07. I validated the suggested approach against main
7227cf8d6e132d33f9b4b8de4d1ae94b8bdb4cf4(ADK 2.10.0), both alone and combined with the current #7323 patch (e05c6cbb203b3eeec51bbbd17de8f018515c130e). The plain-text fix holds up, but I found boundaries that need to be resolved before a PR.The candidate was limited to:
- Add
model_version=self._model_versionto the response returned by__build_full_text_response(). - Filter out events where
event.partialis true before evaluation conversion.
Observed results
Scenario Candidate result One answer split into 1 / 2 / 3 text chunks, no usage Inference counts [1.0, 1.0, 1.0]; final text unchangedTwo completed model turns, same invocation/author/model 2.0Two child agents followed by their parent, through real Runner.run_live()/BaseAgentlifecycles andInMemorySessionService3.0for both streamed and persisted eventsTool-only response with one or two parallel function calls 1.0inference; tool counts1.0/2.0Tool response followed by another model turn 2.0inference; tool call retainedUsage before content, after content but before completion, in the same transport message as content, or on the completion message With #7323: 15.0tokens and1.0inference for all 1 / 2 / 3 chunk variantsTwo child agents plus parent, each reporting usage With #7323: 45.0tokens and3.0inference for both streamed and persisted eventsNon-Live completed-event fixtures, including a tool cycle Counts and usage unchanged Without #7323, the separate usage reports still disappear, as expected from #7321. The current Gemini Live adapter emits usage separately even when usage and content arrive in the same transport message.
Remaining boundaries
- Thought text followed by answer text in one completed turn still counts as
2.0. This happens with both separate transport messages and thought/text parts in the same message. The adapter emits a non-partial response when the thought flag changes, and another for the answer; stamping both makes both eligible model-call events. For the same-message fixture, unmodified main counted1.0, so the candidate changes that result to2.0. - Text followed by a tool call before the same
turn_completealso counts as2.0. The text flush and aggregated function-call response are separate non-partial model-version-bearing events. The function call and text are preserved, but the count remains two. Under a one-call-per-logical-response interpretation, this still overcounts; the intended semantics need confirmation. - A usage-bearing partial
Eventloses its usage when filtered. A constructed fixture containing a partial event with 15 tokens followed by its non-partial final event produces1.0inference but unavailable token usage, including with fix(eval): preserve live usage without extra inference calls #7323. This is a conversion-level fixture, not a claim that the current Gemini Live adapter attaches usage to its partial text events. A shared filter would need to account for this shape or explicitly constrain its scope.
#7323 coordination and validation limits
I ran the existing Live-connection, evaluation-conversion, and efficiency-evaluator test files on Python 3.12.13 /
google-genai2.24.0:- Main: 152 passed.
- Candidate alone: 152 passed.
- Main + fix(eval): preserve live usage without extra inference calls #7323: 166 passed.
- Candidate + fix(eval): preserve live usage without extra inference calls #7323: 162 passed, 4 failed. All four failures are the two-chunk variants of
test_live_usage_preserved_without_extra_model_calls: usage/final-text checks pass, then the existing== len(chunks)assertion expects 2 while the new result is 1.
The additional 35-case diagnostic matrix checks counts, tokens, tool calls, final text, and that conversion does not mutate input events. The combined candidate matches the proposed expectations in 31 cases; the remaining four are the two thought/text shapes, text/tool shape, and partial-usage fixture above. The sub-agent experiment uses real Runner, agent lifecycle, session, Live adapter, and evaluator implementations with a deterministic local transport; it does not use a remote Gemini model or exercise LlmAgent's live model/tool execution loop. I have not run the full Python-version matrix or installed-wheel end-to-end validation for this experimental candidate.
Could the eval owners confirm how the thought/text and text/tool records within one logical model response should count? Non-partial events alone do not appear to identify logical inference boundaries. Once that is agreed, I can keep #7351 focused and align its regression expectations and landing order with #7323.
Minimal local reproduction of the candidate boundaries
Run this on the pinned main snapshot, or that snapshot with #7323 applied. It stamps merged text through a temporary method wrapper and applies the proposed partial filter before conversion; it makes no remote model calls.
"""Run on main, or main + #7323, to reproduce candidate boundaries.""" import asyncio from unittest.mock import patch from google.adk.evaluation._efficiency_evaluators import _InferenceCallCountV1Evaluator from google.adk.evaluation._efficiency_evaluators import _TokenUsageV1Evaluator from google.adk.evaluation.evaluation_generator import EvaluationGenerator from google.adk.events.event import Event from google.adk.models.gemini_llm_connection import GeminiLlmConnection from google.genai import types MODEL = 'gemini-live-test' class LocalTransport: session_id = 'local-test' def __init__(self, messages): self.messages = messages async def receive(self): for message in self.messages: yield message def text(text, thought=False): return types.LiveServerMessage(server_content=types.LiveServerContent( model_turn=types.Content(role='model', parts=[types.Part(text=text, thought=thought)]) )) def finish(): return types.LiveServerMessage(server_content=types.LiveServerContent(turn_complete=True)) def scores(events): # Apply exactly the proposed partial-event filter before conversion. invocations = EvaluationGenerator.convert_events_to_eval_invocations( [event for event in events if not event.partial] ) return ( _InferenceCallCountV1Evaluator().evaluate_invocations(invocations).overall_score, _TokenUsageV1Evaluator().evaluate_invocations(invocations).overall_score, ) async def run(messages): connection = GeminiLlmConnection(LocalTransport(messages), model_version=MODEL) events = [Event(author='user', invocation_id='inv1', content=types.Content( role='user', parts=[types.Part(text='Hi')] ))] async for response in connection.receive(): events.append(Event(author='agent', invocation_id='inv1', **response.model_dump(exclude_none=True))) return scores(events) async def main(): method = '_GeminiLlmConnection__build_full_text_response' original = getattr(GeminiLlmConnection, method) def stamp(connection, *args, **kwargs): # Equivalent to adding model_version=self._model_version to this response. return original(connection, *args, **kwargs).model_copy(update={'model_version': MODEL}) with patch.object(GeminiLlmConnection, method, stamp): assert await run([text('Hel'), text('lo'), finish()]) == (1.0, None) thought = await run([text('Reasoning', True), text('Hello'), finish()]) tool = await run([text('Checking'), types.LiveServerMessage( tool_call=types.LiveServerToolCall(function_calls=[types.FunctionCall( id='tool1', name='lookup', args={'query': 'test'} )]) ), finish()]) print('one turn: thought + answer:', thought) print('one turn: text + tool:', tool) assert thought == (2.0, None) assert tool == (2.0, None) usage = types.GenerateContentResponseUsageMetadata( prompt_token_count=10, candidates_token_count=5, total_token_count=15 ) fixture = [ Event(author='agent', invocation_id='inv1', partial=True, model_version=MODEL, content=types.Content(role='model', parts=[types.Part(text='Hel')]), usage_metadata=usage), Event(author='agent', invocation_id='inv1', partial=False, model_version=MODEL, content=types.Content(role='model', parts=[types.Part(text='Hello')])), ] print('constructed partial Event with usage:', scores(fixture)) assert scores(fixture) == (1.0, None) asyncio.run(main())
Output on both main and main + #7323:
one turn: thought + answer: (2.0, None) one turn: text + tool: (2.0, None) constructed partial Event with usage: (1.0, None)- Add
Great testing @yang0228. Getting the same three results on 2.10.0 so you're right that partial + model_version alone can't mark where a call ends because the Live adapter also flushes non-partial events on thought switches and before tool calls.
On semantics: telemetry counts one inference per model request (record_inference_telemetry in _finalizer.py) so thought+answer or text+tool in one response should be 1. A workaround tried locally: count a call when an author's first model_version/usage event appears and close it on turn_complete or a function_response/user event (Gemini 3.x Live doesn't send turn_complete until the tool result arrives). No partial filter so usage stays and gets summed. That gave 1 for 1/3 chunks, thought+answer, text+tool and your partial-usage fixture and 2 for two real turns and for tool → response → answer.
Could you try that grouping in the conversion/evaluator path against your 35-case matrix including the sub-agent runs? Please test it thoroughly before raising the PR. The 4 failures with #7323 come from its == len(chunks) assertion which encodes the current per-chunk count so let #7323 land first and update that assertion to 1 in your PR.
@i-yliu could you or the eval owners confirm "one model request = one inference call" for Live so yang0228 can go ahead?
- addedlive[Component] This issue is related to live, voice and video chat[Component] This issue is related to live, voice and video chat
on Oct 1, 2026 Thanks @surajksharma07. I implemented the grouping locally in the conversion/evaluator path on main
6be386d8939c4898aa2c0fc8a1f7cb34e7df7904(ADK 2.11.0), with the current #7323 patch (e05c6cbb203b3eeec51bbbd17de8f018515c130e) applied as a prerequisite. All 35 cases from the previous diagnostic matrix now match the expected results.The implementation reconstructs calls from the original event sequence, before conversion drops completion markers. It keeps an active call per author/Live connection, closes it on completion, interruption, a tool response, or a new client user event, and distinguishes ordinary SSE final responses. An optional
Invocation.inference_call_countcarries the result through eval JSON serialization; older eval files and unary-only invocations retain the existing event-based fallback. Merged final text carriesmodel_version, so persisted sessions that omit partial chunks still count correctly. No partial-event filter is used, and standalone usage already merged by #7323 cannot open another call.Scenario Result One answer in 1 / 2 / 3 chunks 1 inference; final text retained Usage surfaced before/after content, combined with content, or on completion 1 inference and 15 tokens Thought + answer, split or in one transport message 1 Text + tool, tool-only, or parallel tools in one response 1; tool calls retained Two genuine model turns, or tool → response → answer 2 Two sub-agents plus parent through real Runner/session lifecycles 3, for streamed and persisted events; 45 tokens when each reports 15 Constructed partial-usage fixture 1 and 15 tokens, including after eval JSON round-trip Non-Live completed calls and tool cycles Existing counts and usage retained Further regression tests exposed three boundaries beyond that matrix: interrupted Gemini 3.x completions can carry grounding metadata; Live and unary calls from the same author must still respect #7323's usage merging; and input audio transcription events can arrive during a model turn. The last case requires preserving
live_session_idon transcription events, so both raw and normalized transcriptions remain distinguishable from a newly submitted user message. These cases are covered, along with interleaved authors, connection changes, non-Live SSE calls, and usage delivered on a later receive after completion. The latter uses the real LLM Live receiver with deterministic transport; it establishes the consumer behavior, not actual Gemini server ordering.Validation
- Related evaluation, adapter, receiver, and Live-flow tests: 245 passed on Python 3.12.13 /
google-genai2.24.0. - Original 35-case matrix: 35/35 against the final source, an independently installed wheel outside the checkout, and each Python version from 3.10 to 3.14. The isolated version environments use
google-genai2.28.0. - Pre-commit checks passed. Mypy on the five changed source files reports the same four pre-existing errors as the baseline, with no new errors.
- Full unit suites, using temporary tox runner overrides because the checkout has no
uv.lock:
Python Passed Failed 3.10.20 17,513 3 3.11.15 17,522 3 3.12.13 17,513 3 3.13.15 17,513 3 3.14.7 17,513 3 The three failures on every version are
test_load_web_page_allows_public_nat64_ip,test_load_web_page_fetches_public_urls_by_pinning_the_resolved_ip, andtest_load_web_page_tries_another_resolved_address_after_connect_error. All three reproduce on the unchanged baseline in each environment. On Python 3.12, isolating the macOS system proxy in the diagnostic process makes that baseline test file pass all 29 tests. The full-suite results above still include those failures, rather than being reported as a clean matrix.The transport is mocked; adapter, receiver, conversion, evaluators, and the relevant Runner/session lifecycles are real. This is not remote Gemini service or release integration validation.
#7323 is still open and the eval-owner confirmation is still pending. I have kept this fix local and have not raised a PR. Once #7323 lands and the semantics are confirmed, I will rebase and include the assertion update from
len(chunks)to1in the #7351 PR, as requested.- Related evaluation, adapter, receiver, and Live-flow tests: 245 passed on Python 3.12.13 /
Really thorough work @yang0228. Checked 2.11.0 and the per-event count in _is_model_call_event() is still there and #7323 is still waiting on review so nothing upstream has changed under you.
Two things eval owners will likely look at closely: the new Invocation.inference_call_count field changes the saved eval-set JSON so please note in the PR why it can't be derived from invocation_events; and the live_session_id gap is real (_transcription_manager.py builds the transcription Event without it) so keeping that as its own small commit would make it easier to review.
Rather than keeping it local suggest opening it as a draft PR stacked on #7323 so reviewers can look at the actual diff, still please test it thoroughly before marking it ready for review.
@i-yliu when you get a chance could you take a look at #7323 and confirm "one model request = one inference call" for Live here?
Opened the dependent draft PR #7427, as suggested. It contains refreshed #7323 and targets
main; the draft description links an incremental comparison that excludes the prerequisite.#7323 is now at
506784c478c5feebe7cc203d35e7b0de9573a549, synced with mainfac77be5d32cf8b07db0162fdd73f5f9155f60bfwithout rewriting its published history. I also fixed two usage-conversion edges found during review: reports attached to nontext parts, and usage from a different Live connection. Its current per-event/chunk inference expectation is preserved; the draft changes that assertion to1separately.The draft has the requested separate transcription commit (
bc5f755886e2a309c895715802faed834b0f2336), followed by the grouping commit (89a65a035e50cdb7b5726c636170c1b59f5cf590). The production receiver preserves transcription connection identity through Runner persistence and eval normalization. Grouping now also handles interleaved same-author connections, foreign connection boundaries, and contentless unary/SSE responses in mixed invocations.The persisted-field rationale and compatibility boundary are explicit in the PR:
invocation_eventslacks the original partial/completion/interruption/connection metadata, so it cannot reconstruct logical request boundaries after projection. Missing/Nonekeeps the old counting fallback. Older eval JSON remains readable by the new implementation, but older strict readers may reject new JSON containing the proposed field.Fresh validation:
- 244 related tests passed; the twelve new review-edge regressions failed before their corresponding fixes and passed afterward.
- 35/35 diagnostic cases on final source, an independently installed wheel outside the checkout, and Python 3.10–3.14.
- Pre-commit passed; type checking introduced no errors beyond the four on unchanged main.
- Full-version results and baseline comparisons are in the draft. They retain the three macOS proxy failures and Python 3.13's additional unraisable recursion failure, also observed on unchanged upstream code; they are not reported as all-green.
The transport is deterministic/local; this does not establish remote Gemini ordering, billing semantics or release integration.
The PR remains draft, pending #7323 landing and eval-owner confirmation of both “one model request = one inference call” and the optional saved-schema field. After the prerequisite is imported, I can rebase away its inherited diff before marking this ready.
Went through #7427 @yang0228. _count_streamed_inference_calls doing the grouping on the raw events before conversion drops partial/turn_complete is the right place for it and the fallback to None when there's no stream key keeps the run_async/adk eval path on the old count. The compatibility note on the new Invocation field reads well. On 2.11.0 _is_model_call_event is still per event so nothing upstream has moved.
One correction on my last comment: TranscriptionManager has no callers on 2.11.0 so putting live_session_id on the transcription events in postprocess_live_flow was the right fix not _transcription_manager.py.
Two small things before it leaves draft. The # ponytail: note in the usage-merge loop looks like a leftover tag so a plain TODO would read better. There's also no test for a typed user message (no transcription, no live_session_id) arriving mid-turn. That path clears active_sessions so one test pinning whether it should count as 2 would make the intent clear.
@i-yliu #7323 is the prerequisite here and still needs a review plus approval for its 2 workflows. Could you also confirm "one model request = one inference call" for Live so this can come out of draft?
@yang0228 @surajksharma07 thanks for bringing up this issue; For live: one inference call = one model generation (one model request), regardless of chunking.
For example:
- tool call → tool response → answer = 2,
- an interrupted generation = 1
- Answer streamed in N text/audio chunks → 1
- Thought + answer, or text + tool call, in one generation → 1
- Tool call → tool response → answer → 2
- Interrupted generation → 1
Please add a test with multiple audio chunks in one generation. I'll take a look at #7323.
- added a commit that references this issue
on Oct 8, 2026 @surajksharma07 @i-yliu Addressed the review follow-up in #7427 at
19b61d22b64fe26f085836d631a33d8499881e7b:- A typed user message with neither transcription nor
live_session_id, arriving between an unfinished generation and the next model response, now has an explicit regression asserting 2 inference calls. - Identical output audio in 1 / 2 / 3 chunks, with and without standalone usage, has six regression cases asserting 1 inference call, retained audio bytes, and 15 tokens when usage is present.
- The usage-merge comment is now a plain TODO in fix(eval): preserve live usage without extra inference calls #7323 at
3517a033be5a6ddf12c4e30f5f179105f48cf5bc; that prerequisite update is merged into fix(eval): count logical Live requests across response segments #7427. This follow-up changes tests and a comment, not production behavior.
The seven new cases pass. Isolated mutations make the typed-message test fail at
1 != 2, and four audio cases fail at2/3 != 1when the evaluator is returned to event-based counting.Fresh validation: 252 related tests passed, configured pre-commit hooks passed, and 7 new cases passed from an independently installed wheel outside the checkout. Full unit suites ran on the currently supported Python 3.11–3.14 versions; every new case passed. Their remaining failures were freshly reproduced on unchanged main
315c3dd2cin the same respective environments, including a complete Python 3.13 baseline run. Exact counts, failure IDs, dependency versions, and reproduction commands are in the updated PR testing plan.The description now records the confirmed one-generation/one-inference semantics and the PR's actual ready-for-review status. #7323 still needs to land first; its inherited diff will be removed afterward. The optional
Invocation.inference_call_countsaved-schema change remains an explicit review question: new readers accept older files, while older strict readers may reject files containing the new field. Could you confirm acceptance of that compatibility boundary when reviewing #7427?Both PRs still need maintainer approval for their workflow runs. This validation uses deterministic local transport and does not claim remote Gemini ordering or billing behavior.
- A typed user message with neither transcription nor
Went through 19b61d2 @yang0228. The typed-message test and the 1/2/3 audio-chunk cases cover what was asked. #7427 still merges cleanly on current main (8e386a9) and the 6 related test files pass except test_send_to_model_caches_only_audio_blobs which fails the same way on plain main so it's not from this PR.
One gap against @i-yliu's list: interrupted generation = 1 is tested but not an interrupted generation followed by a new one which is the usual barge-in. Tried interrupt → transcribed "stop" → new answer locally and the PR gives 2 which is correct so a test pinning that would be worth adding.
v2.11.0 still counts per event so nothing upstream has changed.
@i-yliu when you review #7323 could you also approve the workflow runs on both PRs and confirm whether the optional Invocation.inference_call_count field in saved eval JSON is acceptable? That's the last open question on #7427.
@surajksharma07 @i-yliu The requested barge-in regression is now in #7427 via d485a767. It covers interrupted generation → transcribed “stop” → new answer = 2 inference calls, with both raw and normalized transcription, through the real Live receiver and adapter.
258 related tests passed, along with the configured commit hooks and both new cases against an independently installed wheel. Removing the interruption boundary makes both cases fail as expected. The full suite was not completed. These results are for the test change before the subsequent main merge at 970c239; the test file is unchanged, but I have not rerun tests on that merged head.
#7323 must land first. Saved-eval JSON schema acceptance and maintainer approval for the current head’s workflows remain pending.
Describe the Bug
For a single completed Gemini Live model response,
inference_call_count_v1changes with the number of text chunks delivered by the transport. The same final answer produces counts of 1, 2, or 3 when split into 1, 2, or 3 chunks. Transport chunking should not be interpreted as additional model calls or reasoning steps.This reproducer emits no usage metadata, so it isolates the chunk-counting problem from the standalone usage loss in #7321 / PR #7323.
Steps to Reproduce
Check out the tested main commit and install the evaluation dependencies:
Save the minimal reproduction below as
repro_live_inference_count.py.Run it against the checkout:
Observe counts of
[1.0, 2.0, 3.0]and the failed assertion expecting[1.0, 1.0, 1.0]. All three cases contain oneturn_completeand produce the same final response.Expected Behavior
One completed model response should contribute one inference call regardless of how its text is chunked. Genuine additional model calls, including multiple calls or sub-agents within the same invocation, should still be counted separately.
Observed Behavior / Logs
{"text_chunks": 1, "completed_model_turns": 1, "final_response": "Hello world", "inference_call_count_v1": 1.0} {"text_chunks": 2, "completed_model_turns": 1, "final_response": "Hello world", "inference_call_count_v1": 2.0} {"text_chunks": 3, "completed_model_turns": 1, "final_response": "Hello world", "inference_call_count_v1": 3.0}Environment Details
d312c0ec, verified on 2026-09-30.google-genai: 2.24.0.Model Information
gemini-live-testmodel-version label is supplied to the realGeminiLlmConnectionadapter. No remote model, API key, network model request, or tool call is required.Regression
Unknown; not claiming a regression from a particular release. Reproduced on current main.
Additional Context / Source Trace
The affected path is
GeminiLlmConnection.receive()→Event→EvaluationGenerator.convert_events_to_eval_invocations()→_InferenceCallCountV1Evaluator.model_version, including partial chunks.model_version._is_model_call_event()treats anymodel_versionas a model call; the evaluator counts all such events.PR #7323 preserves the existing multi-chunk inference count while fixing standalone token usage. This issue concerns identifying logical calls across chunks, so it needs a separate focused fix. Deduplicating all events by invocation ID or model version would undercount genuine repeated calls and should be avoided.
Minimal Reproduction Code
Frequency
Always (100% in the deterministic local reproduction; repeated with the same output).
Screenshots / Video
N/A — console reproduction.
Contribution / Assignment Request
I would like to work on this issue. Is anyone already addressing it? Following CONTRIBUTING's request to ask before contributing, could a maintainer confirm the intended counting semantics and assign this issue to @yang0228 if this contribution is welcome?
I propose a focused fix in the shared conversion/evaluation path, with regression coverage for:
Before a PR, I will follow CONTRIBUTING's unit-test, Python-version matrix, pre-commit, wheel, and reproducible end-to-end evidence requirements.