fix(gateway): speech DNS off the async workers; reuse vendor connections for one-shot audio - #13
Merged
Merged
Conversation
…c workers `SelfHostedTTS::new_azure_openai` ran `validate_azure_openai_url`, whose resolve-then-validate step is a blocking `getaddrinfo`, inside the synchronous provider constructor -- on a tokio worker, once per /v1/audio/speech request. Under Kubernetes' `ndots:5` an external name takes ~70 ms to resolve, so the replica's two workers served ~25 speech requests a second at ~0.08 cores, and every other session on the pod waited with them: realtime first-audio p99 went 722 -> 1535 ms under 25 req/s of speech. Transcription already runs the same check through `spawn_blocking` (`transcribe.rs`). Speech now does the same: * `core::net::validate_url_for_ssrf_without_dns` -- every check except the resolution (scheme, blocked hostnames, IP literals). The constructor uses it, so a bad URL is still refused at construction with the same message. * `SelfHostedTTS::connect` runs the full check in `spawn_blocking` before the first dial, so a name that resolves to a private address is still refused. Tests: the DNS-free variant refuses everything the full check does except a name that only resolves privately; a speech endpoint on `ip6-localhost` constructs and is refused at `connect` (fails with the connect-time check removed). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…cription Every /v1/audio/speech and /v1/audio/transcriptions call built a new HTTP client, so each paid DNS + TCP + TLS to the vendor, and speech providers without a shared per-vendor manager also ran ReqManager::warmup per call: 4 unauthenticated HEADs to the vendor before the real POST (measured at the mock: 10 speech calls -> 10 POST + 40 HEAD). At ~4 ms of CPU per request this made CPU the per-replica ceiling (~250-360 speech req/s on 2 CPUs). * Speech: `synthesize_once_standard` now hands the provider a pooled manager -- the shared per-vendor one when it exists, otherwise one per deployment (`CoreState::deployment_tts_req_manager`, keyed by vendor, api_base and timeouts; credentials stay per request). A provider given a manager does not warm up. Long-lived /ws session providers are unchanged. * Transcription: the Azure OpenAI, self-hosted and prerecorded paths use `core::net::shared_http_client`, one pooled client per configuration. * Shared clients are HTTP/1.1 only (`ReqManagerConfig::http1_only`): many 5-10 s requests on one shared HTTP/2 connection queue behind the vendor's concurrent-stream limit. * Permits: `WAAV_TTS_MAX_CONCURRENT_PER_DEPLOYMENT` (default 4096); the ReqManager bound rises from 1000 to 10000 (`MAX_CONCURRENT_REQUESTS`), since one manager now carries a replica's whole load for its deployment. Verified live: 0 HEADs per speech call, 3 pooled vendor connections and no TIME_WAIT after 20 calls. Tests: cache built once and shared per key, HTTP/1.1, permit ceiling; lib tests for the touched modules (2168) and the api_tests / transcribe_batch / voice_http_spans suites pass. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…r-deployment TTS manager `core/state.rs` imports `std::time::Duration` only under `turn-detect`, so the default, dag-routing and noise-filter builds failed on the timeouts `deployment_tts_req_manager` copies from TTSConfig (E0433). Fully qualified. Checked: `cargo check --lib --bins --tests` with no features, dag-routing, noise-filter and dag-routing,turn-ensemble,noise-filter,openapi; changed-module lib tests with default features (117 passed). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Two changes that let one WaaV replica serve the HTTP audio routes at the vendor's own latency, and let a CPU-based HPA scale it:
fix(gateway): resolve the Azure OpenAI speech SSRF check off the async workers (1484e3d)perf(gateway): reuse vendor connections for one-shot speech and transcription (54850c5)Problem (measured on pde-ditto, WaaV 0.1.4, 1 replica, 2 CPU)
A mock vendor behind a public name answered transcription in 10 s and speech in 5 s.
SelfHostedTTS::new_azure_openairanvalidate_azure_openai_url, whose resolve-then-validate step callsgetaddrinfo(core/net.rs,to_socket_addrs). It ran inside the synchronous provider constructor, so on a tokio worker, once per request. Under Kubernetes'ndots:5an external name takes ~70 ms to resolve, so a replica's two workers topped out at ~2 / 0.07 s. The effects:ReqManager::warmupon every call. At the mock, 10 speech calls arrived as 10POSTs plus 40 unauthenticatedHEADs.Changes
Commit 1: DNS off the workers. Transcription already ran this check in
spawn_blocking(handlers/transcribe.rs); speech now does the same.core::net::validate_url_for_ssrf_without_dnsruns every SSRF rule except the resolution: scheme, blocked hostnames, and every IP-literal spelling. The speech constructor uses it, so a bad URL is still refused at construction with the same message.SelfHostedTTS::connectruns the full check inspawn_blockingbefore the first dial, so a name that resolves to a private address is still refused.Commit 2: connection reuse.
synthesize_once_standardhands the provider a pooled manager: the shared per-vendor one when it exists, otherwise one per deployment (CoreState::deployment_tts_req_manager). The per-deployment one is keyed by vendor,api_baseand timeouts, while credentials stay per request. A provider given a manager doesn't warm up. Long-lived/wssession providers are unchanged.core::net::shared_http_client, one pooled client per configuration.ReqManagerConfig::http1_only). Many 5–10 s requests on one shared HTTP/2 connection would queue once the vendor's concurrent-stream limit is reached.WAAV_TTS_MAX_CONCURRENT_PER_DEPLOYMENT(default 4096): permits of the per-deployment manager, bounding transport concurrency on the replica only. Vendor-account concurrency is still each deployment'smax_concurrent.ReqManagerpermit bound raised from 1,000 to 10,000 (MAX_CONCURRENT_REQUESTS), because one manager now carries a replica's whole load for its deployment.Results (same rig; p50 / p99)
POSTs and 0HEADs. After 20 calls: 3 pooled vendor connections, noTIME_WAIT.Testing
New tests:
ip6-localhostconstructs and is refused atconnect. This test fails with the connect-time check removed (verified).Notes
ndots:5. Caching the verdict per host for a short TTL would remove that.ESTABLISHEDsockets on the replica.🤖 Generated with Claude Code