fix(gateway): record each voice analytics attribute once, for what actually ran - #14
Merged
Merged
Conversation
…tually ran The deployment analytics panel (budmetrics VoiceTurnFact / VoiceSessionFact) read wrong or missing values from the voice.turn / voice.session spans: 1. Fallback-served calls exported vendor/model/voice twice. open_turn recorded the PRIMARY's vendor+model (and the plan its voice), a served fallback recorded its own again, and tracing-opentelemetry 0.30 APPENDS a KeyValue for a re-recorded field - so ClickHouse's SpanAttributes['<key>'] paired one hop's vendor with the other's model. The same held for bud.voice.stt.noise_suppression (every attempt of every hop) and the root's gen_ai.provider.name / gen_ai.request.model. VoiceSpans now holds the call's Leg (vendor, EFFECTIVE model, voice, language, denoiser flag) and writes it once when the call ends: the served deployment on success, the primary when every hop failed. The model is the one the vendor was called with (stt.model substitutes; TTS the planned model), not merely the saved endpoint.model. record_vendor is gone. 2. Circuit-breaker fast-fails were recorded as vendor_5xx with no vendor status - blaming a vendor that was never called. They are now error_type=circuit_open (a new member of the closed FRD-021 vocabulary). When the primary's breaker was open and a fallback DID call its vendor and fail, that last real failure's class and vendor status are recorded instead. 3. TTS output length for compressed formats: output_audio_seconds / sample_rate were NULL for mp3/opus/aac/flac. waav_openai_audio::pcm::container_timing demuxes (never decodes): packet durations summed for MP3/ADTS/FLAC/MP4 (the readers' bitrate-estimated n_frames are not trusted; LAME delay/padding subtracted), the last granule position less pre-skip for Ogg. Only bytes that open with one of those containers are read; anything else stays NULL. The rate comes from the container, else ElevenLabs' output_format string - never the synthesis' labelled (defaulted) rate. Consequence: per-second/per-minute priced TTS in compressed formats is now priced. 4. Realtime: a vendor closing its socket normally (1000/1001) ended the session as upstream_error with ERROR status. It is now end_reason=vendor_close, the vendor's code relayed to the client with no error event, not an error. Abnormal closes stay upstream_error. 5. bud.voice.detected_language is normalized to the lowercase BCP-47 primary language subtag (ISO 639-1 where one exists): en for en, eng, english, en-US. Unlisted codes keep their own region-stripped subtag; und/unknown are absent. 6. Google and Azure Speech now set STTResult::vendor_confidence (Google: final results, not proto3's 0.0; Azure: the top NBest entry, detailed format only), so the batch HTTP route - which replays files through these streaming clients - records bud.voice.stt.confidence. AWS Transcribe reports only per-word confidences and is left unset. Tests: exported-span unit tests (one KeyValue per leg key; the append premise pinned), voice_http_spans fallback/breaker/duration/language cases plus a duplicate-key guard in check_root, a realtime vendor-close case, container timing on synthetic MP3/ADTS/FLAC frames and the committed Ogg Opus fixture, language table, Google/Azure confidence. _typos.toml allows four ISO 639-2 codes the language table lists. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
The voice deployment analytics panel reads its numbers from the
voice.turnandvoice.sessionspans WaaV exports. Those spans are the source for budmetrics'VoiceTurnFact/VoiceSessionFacttables. Six of their fields were wrong or missing. This PR records each one correctly, once, for what actually ran.Companion PR: BudEcosystem/bud-runtime#3086, which reads these fields into the redesigned panel. Each side works without the other; see "Compatibility" below.
What was wrong, and what changes
open_turnrecorded the primary's vendor and model, and a served fallback recorded its own on top.SpanAttributes['<key>']then paired one hop's vendor with the other hop's model.bud.voice.stt.noise_suppressionand to the root span'sgen_ai.provider.name/gen_ai.request.model.VoiceSpansholds the call'sLeg(vendor, effective model, voice, language, denoiser flag) and writes it once, when the call ends.endpoint.model.record_vendoris removed.vendor_5xxwith no vendor status. That blamed a vendor that was never called.error_type = circuit_open, a new member of the FRD-021 error vocabulary.output_audio_secondsandsample_ratewere missing.waav_openai_audio::pcm::container_timingdemuxes, never decodes:output_format, never from the synthesis' defaulted label.upstream_error.end_reason = vendor_close: the vendor's close code is relayed to the client, and no error event is sent.upstream_error.bud.voice.detected_languagecame in as the vendor spelled it:en,eng,english,en-US.und/unknownare left absent.STTResult::vendor_confidence:For reviewers
circuit_open(error class) andvendor_close(end reason).circuit_opento a vendor fault, labels it in the UI, and countsvendor_closeas a normal end.circuit_openas its own class and treatsvendor_closeas not failed, so nothing breaks either way._typos.tomlallows four ISO 639-2 codes in the language table (bre,wel,som,yor). Without them typos flagslanguage.rs; I checked both ways.Compatibility and deploy order
Independent. This WaaV change can ship before or after bud-runtime#3086. Every attribute keeps its name, and budmetrics' span-to-fact views read the same keys. Until this ships, the panel's affected cards are sparser: output length for compressed TTS, STT confidence, and correct vendor/model on fallback calls.
Testing
New unit and integration tests:
tests/voice_http_spans.rs: fallback, circuit-breaker, duration and language cases, plus a duplicate-key guard incheck_root.tests/openai_realtime_relay.rs: a vendor-close case.container_timing: synthetic MP3, ADTS and FLAC frames, plus the committed Ogg Opus fixture.Run locally in the bookworm builder (rustc 1.96):
Live on pde-ditto: WaaV built from this branch (
dittops/waav:vpanel-1). The panel's live driver in bud-runtime#3086 (specs/021-voice-analytics/panel_v2_live.py) recomputes every panel number from the raw facts and passed 100/100 and 60/60 checks. Fresh rows showed:The fallback fix is covered by the integration tests only: no live fallback happened during the run, because the primaries stayed healthy.
🤖 Generated with Claude Code