fix: preserve adaptive interruption across tool calls - #6564
Conversation
|
This looks like the right fix — treating the thinking gap as a real speech boundary and restarting overlap inference when playout resumes addresses the root cause at the signal source, where my #6550 only made the stream tolerate the redundant start sentinel. The metrics and remote-session coverage go beyond what mine attempted, too. Closing #6550 in favor of this. If the hermetic stream-level regression tests from it would be useful here (a fake bargein gateway driving |
|
|
||
| def _on_pipeline_reply_done(self, _: asyncio.Task[None]) -> None: | ||
| if not self._speech_q and (not self._current_speech or self._current_speech.done()): | ||
| was_speaking = self._session.agent_state == "speaking" |
There was a problem hiding this comment.
question: what is the difference between session.agent_state and self._agent_speaking?
There was a problem hiding this comment.
(Chiming in from #6550, since this distinction was the heart of #6548 — hope it helps.)
They live at different layers and deliberately diverge mid-turn:
-
session.agent_stateis the session-level lifecycle (listening/thinking/speaking), updated viaAgentActivity._update_agent_state(). A reply that produces tool calls flips it to"thinking"between segments; it is what applications observe throughagent_state_changed. -
AudioRecognition._agent_speakingis a private recognition-side flag meaning "an agent playout window is open for interruption/endpointing purposes". It is set by_on_start_of_agent_speech()(first audio frame of playout) and cleared only inside_on_end_of_agent_speech()— which is intentionally not called when the reply produced tool calls, so it staysTruestraight through the thinking gap. It gates overlap tracking in_on_start_of_speech(), the transcript-ignore window, and_should_ignore_input_audio().
So during a tool-call gap the two disagree by design: agent_state == "thinking" while _agent_speaking == True. was_speaking as written captures "playout was audibly active at reply-done" (session view) rather than "the recognition-side speech window is still open" — which one is wanted here is of course the author's call.
|
from claude code: #6662 section (b) is the same bug with a different trigger: a second agent-speech boundary inside one turn resets the detector stream while an overlap is open. It corrected the false-interruption pause. It did not touch the tool branch: agents/livekit-agents/livekit/agents/voice/agent_activity.py Lines 3443 to 3452 in 59c2fae The tool-response reply is a new Two edits instead:
Then the synthesized Keep the sentinel reorder. It fixes a separate bug, and it is real: the reset currently goes out before the |
c124787 to
4ad72da
Compare
| overlap_started_at = Timestamp() | ||
| overlap_started_at.FromNanoseconds(int(event.overlap_started_at * 1e9)) | ||
|
|
||
| # TODO(AGT-3180): Forward agent_ended when the remote-session protocol supports it. |
There was a problem hiding this comment.
🟡 Inconclusive overlap results are now reported to remote sessions as if the user only backchanneled
Overlaps that end because the agent stopped talking are now published (_on_overlapping_speech at livekit-agents/livekit/agents/voice/remote_session.py:602-606) with no marker that the result is inconclusive, so remote listeners read them as confirmed "user did not interrupt" results.
Impact: Remote dashboards/consumers will count these inconclusive overlaps as backchannels, inflating backchannel statistics for turns where the user may actually have been interrupting.
Why these events start reaching the remote session only now
Before this PR, _on_end_of_agent_speech sent _AgentSpeechEndedSentinel before _on_end_of_overlap_speech, so the detector stream had already reset (_reset_state() clears _overlap_started) and the subsequent _OverlapSpeechEndedSentinel produced no OverlappingSpeechEvent (see livekit-agents/livekit/agents/inference/interruption.py:621-640). The reordering in livekit-agents/livekit/agents/voice/audio_recognition.py:509-514 now emits the event with agent_ended=True, is_interruption=False.
Locally this was compensated for: livekit-agents/livekit/agents/inference/interruption.py:670 no longer counts agent_ended events as backchannels, and AudioRecognition._on_overlap_speech_event skips the backchannel latch for them (livekit-agents/livekit/agents/voice/audio_recognition.py:1480). The remote forwarding path has no such compensation — it only forwards is_interruption, detection_delay, detected_at, overlap_started_at. Until the protocol carries agent_ended (the TODO), an option is to skip forwarding agent_ended events rather than forwarding them as non-interruptions.
Was this helpful? React with 👍 or 👎 to provide feedback.
There was a problem hiding this comment.
Verified this against the branch head (217d25f) — the finding is real, and it is newly reachable behavior rather than a pre-existing gap:
- Before this PR,
_on_end_of_agent_speechpushed_AgentSpeechEndedSentinelbefore_on_end_of_overlap_speech, so_reset_state()had already cleared_overlap_startedand the detector's_OverlapSpeechEndedSentinelarm produced no event. With the reordering (audio_recognition.py:509-514), the detector now emitsOverlappingSpeechEvent(is_interruption=False, agent_ended=True)(interruption.py:621-641) — an event that simply never fired before. AgentActivity._on_overlap_speech_endedforwards every event to the session unconditionally (agent_activity.py:1887-1889), so the remote forwarder sees them all.- Both local consumers were taught the distinction (
interruption.py:670excludesagent_endedfromnum_backchannels;audio_recognition.py:1480skips the backchannel latch), but the wire message cannot express it:OverlappingSpeechin livekit/protocol (protobufs/agent/livekit_agent_session.proto) carries onlyis_interruption/overlap_started_at/detection_delay/detected_at. A remote consumer therefore has no choice but to read these as confirmed backchannels.
Since the proto change lives in livekit/protocol (an optional bool agent_ended = 5 would be backward-compatible there), the self-contained interim fix is to skip forwarding inconclusive verdicts, in line with the existing TODO:
| # TODO(AGT-3180): Forward agent_ended when the remote-session protocol supports it. | |
| # TODO(AGT-3180): Forward agent_ended when the remote-session protocol supports it. | |
| if event.agent_ended: | |
| # an inconclusive verdict is indistinguishable from a confirmed backchannel | |
| # on the wire; skip it until the protocol can carry the marker | |
| return |
Suppressing the event entirely matches pre-PR remote behavior (these overlaps produced no event at all), so nothing downstream regresses in the meantime, and the TODO keeps the protocol follow-up visible.
217d25f to
76e29fd
Compare
76e29fd to
729860b
Compare
Bug: Adaptive interruption could silently drop a user overlap when agent playout resumed after a tool call because the interruption stream received mismatched speech boundaries.
Treat thinking gaps as real speech boundaries and restart overlap inference when playout resumes mid-utterance. Agent-ended overlaps remain inconclusive and no longer count as backchannels; remote forwarding is noted until the protocol carries that marker.
Fixes #6548. Fixes #6548.