Skip to content

fix(gateway): record each voice analytics attribute once, for what actually ran - #14

Merged
dittops merged 1 commit into
mainfrom
fix/voice-analytics-recording
Sep 29, 2026
Merged

dittops merged 1 commit into
mainfrom
fix/voice-analytics-recording

Conversation

@dittops

@dittops dittops commented Sep 29, 2026

Copy link
Copy Markdown
Member

Summary

The voice deployment analytics panel reads its numbers from the voice.turn and voice.session spans WaaV exports. Those spans are the source for budmetrics' VoiceTurnFact / VoiceSessionFact tables. Six of their fields were wrong or missing. This PR records each one correctly, once, for what actually ran.

Companion PR: BudEcosystem/bud-runtime#3086, which reads these fields into the redesigned panel. Each side works without the other; see "Compatibility" below.

What was wrong, and what changes

  1. A fallback-served call exported vendor, model and voice twice.
    • open_turn recorded the primary's vendor and model, and a served fallback recorded its own on top.
    • tracing-opentelemetry 0.30 appends a KeyValue for a re-recorded field. ClickHouse's SpanAttributes['<key>'] then paired one hop's vendor with the other hop's model.
    • The same happened to bud.voice.stt.noise_suppression and to the root span's gen_ai.provider.name / gen_ai.request.model.
    • Change: VoiceSpans holds the call's Leg (vendor, effective model, voice, language, denoiser flag) and writes it once, when the call ends.
      • On success it records the deployment that served the call; when every hop failed, the primary.
      • The model is the one the vendor was actually called with, not just the saved endpoint.model.
      • record_vendor is removed.
  2. Circuit-breaker fast-fails were recorded as vendor_5xx with no vendor status. That blamed a vendor that was never called.
    • They are now error_type = circuit_open, a new member of the FRD-021 error vocabulary.
    • When a fallback did call its vendor and fail, that last real failure's class and status are recorded instead.
  3. TTS output length was NULL for compressed formats (mp3, opus, aac, flac), so output_audio_seconds and sample_rate were missing.
    • waav_openai_audio::pcm::container_timing demuxes, never decodes:
      • MP3, ADTS, FLAC and MP4: sums packet durations, subtracting LAME delay and padding. The readers' bitrate-estimated frame counts are not trusted.
      • Ogg: last granule position minus pre-skip.
    • Only bytes that open with one of those containers are read; anything else stays NULL.
    • The sample rate comes from the container, else from ElevenLabs' output_format, never from the synthesis' defaulted label.
  4. Realtime: a vendor closing its socket normally (1000/1001) ended the session as upstream_error.
    • It is now end_reason = vendor_close: the vendor's close code is relayed to the client, and no error event is sent.
    • Abnormal closes stay upstream_error.
  5. bud.voice.detected_language came in as the vendor spelled it: en, eng, english, en-US.
    • It is now the lowercase BCP-47 primary subtag, ISO 639-1 where one exists.
    • Unlisted codes keep their own region-stripped subtag; und / unknown are left absent.
  6. Google and Azure Speech never reported a confidence, so batch transcription through them recorded none.
    • Both now set STTResult::vendor_confidence:
      • Google: final results only, so proto3's default 0.0 is not read as a score.
      • Azure: the top NBest entry, detailed format only.
    • AWS Transcribe reports only per-word confidences and is left unset.

For reviewers

  • Billing change (item 3): TTS priced per second or per minute in a compressed format now gets a cost. Its output length used to be NULL, so those calls were recorded as unpriced. Character- and request-priced TTS is unchanged.
  • New vocabulary: circuit_open (error class) and vendor_close (end reason).
    • bud-runtime#3086 maps circuit_open to a vendor fault, labels it in the UI, and counts vendor_close as a normal end.
    • An older budmetrics shows circuit_open as its own class and treats vendor_close as not failed, so nothing breaks either way.
  • _typos.toml allows four ISO 639-2 codes in the language table (bre, wel, som, yor). Without them typos flags language.rs; I checked both ways.
  • No config, env var, or dependency changes.

Compatibility and deploy order

Independent. This WaaV change can ship before or after bud-runtime#3086. Every attribute keeps its name, and budmetrics' span-to-fact views read the same keys. Until this ships, the panel's affected cards are sparser: output length for compressed TTS, STT confidence, and correct vendor/model on fallback calls.

Testing

New unit and integration tests:

  • Exported-span unit tests: one KeyValue per leg key, plus a test that pins the append behaviour this depends on.
  • tests/voice_http_spans.rs: fallback, circuit-breaker, duration and language cases, plus a duplicate-key guard in check_root.
  • tests/openai_realtime_relay.rs: a vendor-close case.
  • container_timing: synthetic MP3, ADTS and FLAC frames, plus the committed Ogg Opus fixture.
  • The language table, and Google / Azure confidence.

Run locally in the bookworm builder (rustc 1.96):

cargo fmt --all --check                                                          # clean
cargo clippy --all-targets --features dag-routing,turn-ensemble,noise-filter,openapi
# no warning in any file this PR changes. The one warning in the tree is
# `manual implementation of Option::zip` in stt/iflytek/client.rs, which this PR does not touch and
# which main's CI clippy (stable) does not raise.
typos <changed files>                                                            # clean

Live on pde-ditto: WaaV built from this branch (dittops/waav:vpanel-1). The panel's live driver in bud-runtime#3086 (specs/021-voice-analytics/panel_v2_live.py) recomputes every panel number from the raw facts and passed 100/100 and 60/60 checks. Fresh rows showed:

  • mp3 and opus output seconds and sample rates recorded;
  • a vendor status on rejected calls;
  • no duplicate span keys.

The fallback fix is covered by the integration tests only: no live fallback happened during the run, because the primaries stayed healthy.

🤖 Generated with Claude Code

…tually ran

The deployment analytics panel (budmetrics VoiceTurnFact / VoiceSessionFact) read wrong or
missing values from the voice.turn / voice.session spans:

1. Fallback-served calls exported vendor/model/voice twice. open_turn recorded the PRIMARY's
   vendor+model (and the plan its voice), a served fallback recorded its own again, and
   tracing-opentelemetry 0.30 APPENDS a KeyValue for a re-recorded field - so ClickHouse's
   SpanAttributes['<key>'] paired one hop's vendor with the other's model. The same held for
   bud.voice.stt.noise_suppression (every attempt of every hop) and the root's
   gen_ai.provider.name / gen_ai.request.model. VoiceSpans now holds the call's Leg (vendor,
   EFFECTIVE model, voice, language, denoiser flag) and writes it once when the call ends: the
   served deployment on success, the primary when every hop failed. The model is the one the
   vendor was called with (stt.model substitutes; TTS the planned model), not merely the saved
   endpoint.model. record_vendor is gone.

2. Circuit-breaker fast-fails were recorded as vendor_5xx with no vendor status - blaming a
   vendor that was never called. They are now error_type=circuit_open (a new member of the
   closed FRD-021 vocabulary). When the primary's breaker was open and a fallback DID call its
   vendor and fail, that last real failure's class and vendor status are recorded instead.

3. TTS output length for compressed formats: output_audio_seconds / sample_rate were NULL for
   mp3/opus/aac/flac. waav_openai_audio::pcm::container_timing demuxes (never decodes): packet
   durations summed for MP3/ADTS/FLAC/MP4 (the readers' bitrate-estimated n_frames are not
   trusted; LAME delay/padding subtracted), the last granule position less pre-skip for Ogg.
   Only bytes that open with one of those containers are read; anything else stays NULL. The
   rate comes from the container, else ElevenLabs' output_format string - never the synthesis'
   labelled (defaulted) rate. Consequence: per-second/per-minute priced TTS in compressed
   formats is now priced.

4. Realtime: a vendor closing its socket normally (1000/1001) ended the session as
   upstream_error with ERROR status. It is now end_reason=vendor_close, the vendor's code
   relayed to the client with no error event, not an error. Abnormal closes stay upstream_error.

5. bud.voice.detected_language is normalized to the lowercase BCP-47 primary language subtag
   (ISO 639-1 where one exists): en for en, eng, english, en-US. Unlisted codes keep their own
   region-stripped subtag; und/unknown are absent.

6. Google and Azure Speech now set STTResult::vendor_confidence (Google: final results, not
   proto3's 0.0; Azure: the top NBest entry, detailed format only), so the batch HTTP route -
   which replays files through these streaming clients - records bud.voice.stt.confidence.
   AWS Transcribe reports only per-word confidences and is left unset.

Tests: exported-span unit tests (one KeyValue per leg key; the append premise pinned),
voice_http_spans fallback/breaker/duration/language cases plus a duplicate-key guard in
check_root, a realtime vendor-close case, container timing on synthetic MP3/ADTS/FLAC frames
and the committed Ogg Opus fixture, language table, Google/Azure confidence. _typos.toml
allows four ISO 639-2 codes the language table lists.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@dittops
dittops merged commit 0cb9f5d into main Sep 29, 2026
16 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant