Skip to content

fix(gateway): speech DNS off the async workers; reuse vendor connections for one-shot audio - #13

Merged
dittops merged 3 commits into
mainfrom
fix/tts-ssrf-dns-off-runtime
Sep 29, 2026
Merged

dittops merged 3 commits into
mainfrom
fix/tts-ssrf-dns-off-runtime

Conversation

@dittops

@dittops dittops commented Sep 29, 2026 •

Copy link
Copy Markdown
Member

Summary

Two changes that let one WaaV replica serve the HTTP audio routes at the vendor's own latency, and let a CPU-based HPA scale it:

  1. fix(gateway): resolve the Azure OpenAI speech SSRF check off the async workers (1484e3d)
  2. perf(gateway): reuse vendor connections for one-shot speech and transcription (54850c5)

Problem (measured on pde-ditto, WaaV 0.1.4, 1 replica, 2 CPU)

A mock vendor behind a public name answered transcription in 10 s and speech in 5 s.

  • Speech had a hard ceiling of ~25 req/s per replica, with CPU idle (0.08 cores). SelfHostedTTS::new_azure_openai ran validate_azure_openai_url, whose resolve-then-validate step calls getaddrinfo (core/net.rs, to_socket_addrs). It ran inside the synchronous provider constructor, so on a tokio worker, once per request. Under Kubernetes' ndots:5 an external name takes ~70 ms to resolve, so a replica's two workers topped out at ~2 / 0.07 s. The effects:
    • At 200 speech requests in flight: p50 6.9 s and p99 9.4 s, against the vendor's 5.3 s.
    • At 800 in flight every request failed (503s and 429s).
    • At 1,600 the liveness probe timed out and the kubelet restarted the container.
    • It stalled every session on the replica: 100 realtime sessions saw first-audio p99 go from 0.72 s to 4.87 s while speech ran alongside.
  • Every audio request built a new HTTP client.
    • Each call paid DNS, TCP and TLS to the vendor.
    • Speech providers without a shared per-vendor manager also ran ReqManager::warmup on every call. At the mock, 10 speech calls arrived as 10 POSTs plus 40 unauthenticated HEADs.
    • ~4 ms of CPU per speech request made CPU the next per-replica ceiling (~360 req/s, throttled).

Changes

Commit 1: DNS off the workers. Transcription already ran this check in spawn_blocking (handlers/transcribe.rs); speech now does the same.

  • core::net::validate_url_for_ssrf_without_dns runs every SSRF rule except the resolution: scheme, blocked hostnames, and every IP-literal spelling. The speech constructor uses it, so a bad URL is still refused at construction with the same message.
  • SelfHostedTTS::connect runs the full check in spawn_blocking before the first dial, so a name that resolves to a private address is still refused.
  • The SSRF check is kept, not removed.

Commit 2: connection reuse.

  • Speech: synthesize_once_standard hands the provider a pooled manager: the shared per-vendor one when it exists, otherwise one per deployment (CoreState::deployment_tts_req_manager). The per-deployment one is keyed by vendor, api_base and timeouts, while credentials stay per request. A provider given a manager doesn't warm up. Long-lived /ws session providers are unchanged.
  • Transcription: the Azure OpenAI, self-hosted and prerecorded paths use core::net::shared_http_client, one pooled client per configuration.
  • HTTP/1.1 for the shared pools (ReqManagerConfig::http1_only). Many 5–10 s requests on one shared HTTP/2 connection would queue once the vendor's concurrent-stream limit is reached.
  • New env var WAAV_TTS_MAX_CONCURRENT_PER_DEPLOYMENT (default 4096): permits of the per-deployment manager, bounding transport concurrency on the replica only. Vendor-account concurrency is still each deployment's max_concurrent.
  • ReqManager permit bound raised from 1,000 to 10,000 (MAX_CONCURRENT_REQUESTS), because one manager now carries a replica's whole load for its deployment.

Results (same rig; p50 / p99)

test in flight direct to vendor 0.1.4 + commit 1 + commit 2
speech 200 5.27 / 5.80 s 6.93 / 9.39 s, capped at 23.5 req/s 5.39 / 5.92 s 5.36 / 5.87 s
speech 800 5.29 / 5.80 s all failed 5.39 / 5.93 s 5.36 / 5.88 s
speech 1,600 5.29 / 5.80 s container restarted 5.41 / 5.95 s, 1.04 cores 5.37 / 5.89 s, 0.34 cores
speech 3,200 5.28 / 5.79 s, 508 req/s nothing completed 7.24 / 8.49 s, 364 req/s, throttled 5.36 / 5.90 s, 505 req/s, 0.61 cores
transcription 3,200 10.03 / 11.09 s 10.17 / 11.31 s, 0.44 cores 10.18 / 11.28 s 10.09 / 11.15 s, 0.24 cores
realtime first audio, 100 sessions + speech in background — — 0.82 / 4.87 s 0.62 / 0.72 s 0.62 / 0.72 s
  • CPU per request: speech ~4 ms → ~1.2 ms, transcription ~2 ms → ~1 ms.
  • With the matching chart HPA (CPU 70%, 2–10 replicas; BudEcosystem/bud-runtime#3081), 10,000 requests in flight scaled 2 → 3 → 5 replicas at ~1,410 req/s. Latency stayed flat and there were 0 errors in 687,835 requests.
  • Checked on the running build: 10 speech calls → 10 POSTs and 0 HEADs. After 20 calls: 3 pooled vendor connections, no TIME_WAIT.

Testing

# bookworm builder, --no-default-features --features dag-routing,turn-ensemble,noise-filter,openapi
cargo test --lib -- utils::req_manager:: core::net:: core::stt::prerecorded:: handlers::transcribe:: \
  handlers::speak:: handlers::openai_audio:: core::tts:: core::state::      # 2168 passed
cargo test --test voice_http_spans --test transcribe_batch --test api_tests # 25 + 5 + 5 passed (6 ignored)
cargo fmt --check                                                           # changed files clean
cargo clippy --lib --tests --message-format short                           # no findings in changed files

New tests:

  • The DNS-free validator refuses everything the full check does, except a name that only resolves privately.
  • A speech endpoint on ip6-localhost constructs and is refused at connect. This test fails with the connect-time check removed (verified).
  • A per-deployment manager is built once and shared, is HTTP/1.1, and keeps its permit count.
  • A shared HTTP client is built once per key.
  • The permit ceiling is enforced.

Notes

  • No secrets, IAM or Dapr changes. One new optional env var, listed above.
  • Not addressed here:
    • Per-vendor shared managers (ElevenLabs, Deepgram and the other vendors built into the startup map) keep their 64-permit pool per replica over HTTP/2. At 5 s per request that caps near ~13 req/s per vendor per replica. The test only exercised Azure OpenAI, so this is untested.
    • CoreDNS load: the connect-time SSRF check still resolves DNS on every request. It no longer blocks, but at 1,400 req/s it drove CoreDNS to ~470m under ndots:5. Caching the verdict per host for a short TTL would remove that.
    • Idle client connections: WaaV sets no idle timeout on client connections. A client pod that vanished without closing left ~800 dead ESTABLISHED sockets on the replica.

🤖 Generated with Claude Code

dittops and others added 3 commits September 29, 2026 11:10
…c workers

`SelfHostedTTS::new_azure_openai` ran `validate_azure_openai_url`, whose
resolve-then-validate step is a blocking `getaddrinfo`, inside the
synchronous provider constructor -- on a tokio worker, once per
/v1/audio/speech request. Under Kubernetes' `ndots:5` an external name takes
~70 ms to resolve, so the replica's two workers served ~25 speech requests a
second at ~0.08 cores, and every other session on the pod waited with them:
realtime first-audio p99 went 722 -> 1535 ms under 25 req/s of speech.

Transcription already runs the same check through `spawn_blocking`
(`transcribe.rs`). Speech now does the same:

* `core::net::validate_url_for_ssrf_without_dns` -- every check except the
  resolution (scheme, blocked hostnames, IP literals). The constructor uses
  it, so a bad URL is still refused at construction with the same message.
* `SelfHostedTTS::connect` runs the full check in `spawn_blocking` before the
  first dial, so a name that resolves to a private address is still refused.

Tests: the DNS-free variant refuses everything the full check does except a
name that only resolves privately; a speech endpoint on `ip6-localhost`
constructs and is refused at `connect` (fails with the connect-time check
removed).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…cription

Every /v1/audio/speech and /v1/audio/transcriptions call built a new HTTP
client, so each paid DNS + TCP + TLS to the vendor, and speech providers
without a shared per-vendor manager also ran ReqManager::warmup per call:
4 unauthenticated HEADs to the vendor before the real POST (measured at the
mock: 10 speech calls -> 10 POST + 40 HEAD). At ~4 ms of CPU per request this
made CPU the per-replica ceiling (~250-360 speech req/s on 2 CPUs).

* Speech: `synthesize_once_standard` now hands the provider a pooled
  manager -- the shared per-vendor one when it exists, otherwise one per
  deployment (`CoreState::deployment_tts_req_manager`, keyed by vendor,
  api_base and timeouts; credentials stay per request). A provider given a
  manager does not warm up. Long-lived /ws session providers are unchanged.
* Transcription: the Azure OpenAI, self-hosted and prerecorded paths use
  `core::net::shared_http_client`, one pooled client per configuration.
* Shared clients are HTTP/1.1 only (`ReqManagerConfig::http1_only`): many
  5-10 s requests on one shared HTTP/2 connection queue behind the vendor's
  concurrent-stream limit.
* Permits: `WAAV_TTS_MAX_CONCURRENT_PER_DEPLOYMENT` (default 4096); the
  ReqManager bound rises from 1000 to 10000 (`MAX_CONCURRENT_REQUESTS`), since
  one manager now carries a replica's whole load for its deployment.

Verified live: 0 HEADs per speech call, 3 pooled vendor connections and no
TIME_WAIT after 20 calls. Tests: cache built once and shared per key,
HTTP/1.1, permit ceiling; lib tests for the touched modules (2168) and the
api_tests / transcribe_batch / voice_http_spans suites pass.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…r-deployment TTS manager

`core/state.rs` imports `std::time::Duration` only under `turn-detect`, so the
default, dag-routing and noise-filter builds failed on the timeouts
`deployment_tts_req_manager` copies from TTSConfig (E0433). Fully qualified.

Checked: `cargo check --lib --bins --tests` with no features, dag-routing,
noise-filter and dag-routing,turn-ensemble,noise-filter,openapi; changed-module
lib tests with default features (117 passed).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@dittops
dittops merged commit e7b76c5 into main Sep 29, 2026
16 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant