-
Notifications
You must be signed in to change notification settings - Fork 1.9k
test: benchmark safe server-side audio level correction #2240
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Changes from all commits
0fb9dc0
87b173d
f4b932f
7913e98
5cfb16a
f5119a8
cbebfe2
23eaa8f
48f8dd8
2a5bade
ee89503
File filter
Filter by extension
Conversations
Jump to
Diff view
Diff view
There are no files selected for viewing
| Original file line number | Diff line number | Diff line change |
|---|---|---|
|
|
@@ -13,17 +13,43 @@ on: | |
| - main | ||
| paths: | ||
| - "apps/media-server/**" | ||
| - "scripts/benchmark-instant-audio.py" | ||
| - "scripts/benchmark-audio-intelligibility.py" | ||
| - "scripts/test-audio-intelligibility.py" | ||
| - ".github/workflows/docker-build-media-server.yml" | ||
| pull_request: | ||
| paths: | ||
| - "apps/media-server/**" | ||
| - "scripts/benchmark-instant-audio.py" | ||
| - "scripts/benchmark-audio-intelligibility.py" | ||
| - "scripts/test-audio-intelligibility.py" | ||
| - ".github/workflows/docker-build-media-server.yml" | ||
|
|
||
| permissions: {} | ||
|
|
||
| concurrency: | ||
| group: media-server-${{ github.head_ref || github.ref_name }}-${{ inputs.tag || 'latest' }} | ||
| cancel-in-progress: true | ||
|
|
||
| jobs: | ||
| audio-metrics: | ||
| name: Audio intelligibility alignment | ||
| runs-on: ubuntu-24.04 | ||
| timeout-minutes: 5 | ||
| permissions: | ||
| contents: read | ||
| steps: | ||
| - uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 | ||
| with: | ||
| persist-credentials: false | ||
| - name: Verify aligned audio metrics | ||
| env: | ||
| PYTHONDONTWRITEBYTECODE: "1" | ||
| run: | | ||
| python3 -m venv "$RUNNER_TEMP/audio-metrics" | ||
| "$RUNNER_TEMP/audio-metrics/bin/python" -m pip install --disable-pip-version-check --no-input --only-binary=:all: numpy==2.4.1 scipy==1.18.1 pystoi==0.4.1 | ||
|
Contributor
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more.
The new alignment job applies Prompt To Fix With AIThis is a comment left during a code review.
Path: .github/workflows/docker-build-media-server.yml
Line: 50
Comment:
**Binary-Only Install Rejects Pystoi**
The new alignment job applies `--only-binary=:all:` to `pystoi==0.4.1`, but that version is distributed without a binary wheel. Pip therefore rejects the available source distribution, so this required CI job fails before the alignment tests can run.
---
For each issue above, determine whether it is valid and should be fixed. If so, fix it directly. |
||
| "$RUNNER_TEMP/audio-metrics/bin/python" scripts/test-audio-intelligibility.py | ||
|
|
||
| build: | ||
| name: Build Docker Image (${{ matrix.platform }}) | ||
| runs-on: ${{ matrix.runner }} | ||
|
|
@@ -77,6 +103,11 @@ jobs: | |
| env: | ||
| MEDIA_IMAGE: ${{ github.event_name == 'pull_request' && 'cap-media-server:verification' || format('ghcr.io/{0}/cap-media-server@{1}', env.REPOSITORY_OWNER, steps.build.outputs.digest) }} | ||
| run: | | ||
| docker run --rm --network none --entrypoint bun "$MEDIA_IMAGE" test \ | ||
| src/__tests__/lib/audio-quality-policy.test.ts \ | ||
| src/__tests__/lib/audio-quality.integration.test.ts \ | ||
| src/__tests__/lib/audio-quality-formats.integration.test.ts \ | ||
| src/__tests__/lib/audio-quality-benchmark.test.ts | ||
| docker run --rm --network none --entrypoint bun "$MEDIA_IMAGE" test \ | ||
| src/__tests__/lib/recording-verification.integration.test.ts \ | ||
| src/__tests__/lib/job-manager.test.ts | ||
|
|
||
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,85 @@ | ||
| { | ||
| "baseline": { | ||
| "count": 60, | ||
| "hours": 4.463996666666667, | ||
| "medianLufs": -29.205, | ||
| "belowMinus24": 40, | ||
| "peakAboveZero": 9, | ||
| "alreadyLoud": 9 | ||
| }, | ||
| "voice": { | ||
| "count": 60, | ||
| "passing": 42, | ||
| "rejected": 1, | ||
| "unchanged": 17, | ||
| "errors": [], | ||
| "skipReasons": { | ||
| "unsafe-levels": 12, | ||
| "already-loud": 5 | ||
| }, | ||
| "medianInputLufs": -31.744999999999997, | ||
| "medianOutputLufs": -16.84, | ||
| "medianGain": 14.495, | ||
| "maxTruePeak": -1.34, | ||
| "maxDurationDeltaMs": 20.999999999958163, | ||
| "sampleCountsExact": true, | ||
| "medianRealtimeFactor": 0.08116900157594112 | ||
| }, | ||
| "levels": { | ||
| "count": 60, | ||
| "passing": 40, | ||
| "rejected": 0, | ||
| "unchanged": 20, | ||
| "errors": [], | ||
| "skipReasons": { | ||
| "unsafe-levels": 12, | ||
| "already-loud": 5, | ||
| "insufficient-headroom": 3 | ||
| }, | ||
| "medianInputLufs": -32.07, | ||
| "medianOutputLufs": -23.08, | ||
| "medianGain": 10.64, | ||
| "maxTruePeak": -1.95, | ||
| "maxDurationDeltaMs": 20.999999999958163, | ||
| "sampleCountsExact": true, | ||
| "medianRealtimeFactor": 0.09984551719406805 | ||
| }, | ||
| "holdoutVoice": { | ||
| "count": 21, | ||
| "passing": 15, | ||
| "rejected": 1, | ||
| "unchanged": 5, | ||
| "errors": [], | ||
| "skipReasons": { | ||
| "unsafe-levels": 2, | ||
| "already-loud": 3 | ||
| }, | ||
| "medianInputLufs": -30.92, | ||
| "medianOutputLufs": -16.78, | ||
| "medianGain": 13.339999999999998, | ||
| "maxTruePeak": -1.34, | ||
| "maxDurationDeltaMs": 20.999999999958163, | ||
| "sampleCountsExact": true, | ||
| "medianRealtimeFactor": 0.08261003124306479 | ||
| }, | ||
| "codeHash": "40ef7500ef8643decd898509879bfe0293ced1f5d124f8a93535f9080f1bf609", | ||
| "scope": "Offline public/transcribed cohort; passing technical gates does not establish perceptual quality or production eligibility.", | ||
| "productionEnabled": false, | ||
| "reviewValidation": { | ||
| "workerCodeHash": "d1dbc4b3be3ff1b833b51a5a87b0933b2b7c82eb0816e7eba690fee12c7a7cf1", | ||
| "existingCohortRetested": 60, | ||
| "existingOutputHashesIdentical": true, | ||
| "additionalProductionFiles": 16, | ||
| "additionalProductionFilesUnchanged": 16, | ||
| "additionalProductionOriginalHashesPreserved": true, | ||
| "localTests": 46, | ||
| "localAssertions": 148, | ||
| "alignmentUnitTests": 2, | ||
| "alignedIntelligibility": { | ||
| "count": 39, | ||
| "medianStoiDelta": -0.00025973077349839, | ||
| "worstStoiDelta": -0.005509516622390076, | ||
| "belowMinusPointZeroOne": 0 | ||
| } | ||
| } | ||
| } |
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,155 @@ | ||
| # Instant audio quality experiment | ||
|
richiemcilroy marked this conversation as resolved.
|
||
|
|
||
| This is an offline, shadow-only experiment. No route, recording finalizer, player, | ||
| export, upload, desktop capture path, or production flag imports the worker. | ||
| `mode: "off"` returns before filesystem access. There is no publishing mode. | ||
|
|
||
| The proposed first rollout is bounded, constant level correction, after production | ||
| validation. EQ and denoising remain experimental because consistent perceptual | ||
| improvement has not been established. This PR does not enable either profile. | ||
|
|
||
| ## Evidence and limits | ||
|
|
||
| The September 7–8, 2026 study measured 60 additional public, unprotected Instant | ||
| recordings with completed transcripts: 20 from each of September 5, 6, and 7, | ||
| all from different owners and separate from the initial 12-recording study. | ||
| The expanded cohort contains 4.464 hours of audio. Median playback-compensated | ||
| loudness is -29.205 LUFS; 40/60 recordings are below -24 LUFS, 9 already exceed | ||
| -18 LUFS, and 9 have true peaks above 0 dBTP. | ||
|
|
||
| A fixed 39-recording tuning / 21-recording holdout split preceded processing. | ||
| The final voice policy passed technical gates on 42/60, left 17 unchanged, and | ||
| rejected one short holdout clip for excessive gain. Among the 42 passing clips, | ||
| median loudness moved from -31.745 to -16.84 LUFS, median gain was 14.495 dB, | ||
| and the highest encoded true peak was -1.34 dBTP. All decoded sample counts were | ||
| preserved; container duration changes were at most 21 ms. The reserved holdout | ||
| alone had 15 passes, five unchanged, and that one rejected candidate. | ||
| Constant-gain processing passed 40/60 and left 20 unchanged, with no rejected | ||
| outputs. These counts measure technical eligibility, not listening preference. | ||
|
|
||
| Five initial policies produced 195 comparisons. Naive dynamic normalization, | ||
| EQ/compression, and denoising each changed container duration by more than 25 ms | ||
| on 29/39 recordings. Aggressive processing also damaged synthetic intelligibility | ||
| scores. These failures remain in the local evidence; they are not shipping presets. | ||
|
|
||
| The adjusted policy uses a 60 Hz high-pass, -0.75 dB at 250 Hz, +0.75 dB at 2.5 kHz, | ||
| 6 dB adaptive FFT denoising, and loudness normalization. It has no extra compressor. | ||
| Input and output gates bound gain and peak level. Separate constant-gain processing | ||
| is available for content whose suitability for voice processing is unknown. | ||
|
|
||
| The exact final worker was tested on 39 controlled cases at 48 kHz: three reference | ||
| voice excerpts, four additive noise types, and three SNRs, plus the unmodified | ||
| references. Median STOI change was -0.000260, worst -0.005510; none exceeded the | ||
| chosen -0.01 regression tolerance. Six inputs were conservatively left unchanged. | ||
| These are relative tests against existing recordings, not clean studio ground | ||
| truth, subjective quality ratings, or a matched Loom comparison. The calibration | ||
| run with 12 dB denoising exceeded that tolerance in four cases, motivating 6 dB. | ||
|
|
||
| Volume-matched RMS in uncaptioned intervals changed by a median +0.096 dB across | ||
| 41 passing clips, with a maximum increase of 6.841 dB. Twelve voice candidates | ||
| changed LRA by more than two LU. Those observations require listening review for | ||
| background noise swelling and altered dynamics before enabling voice processing. | ||
|
|
||
| The reviewed intelligibility scorer aligns reference, noisy input, and processed | ||
| audio to the same overlapping sample interval before computing STOI and SI-SDR. | ||
| It retains unaligned scores and the measured lag separately; aligning a score does | ||
| not waive timing gates. The 39 exact-worker cases were rescored from their original | ||
| artifacts after this correction, with zero cases below the -0.01 tolerance. | ||
|
|
||
| Full-recording LUFS and caption-aligned RMS answer different questions. Caption | ||
| intervals approximate speech activity; uncaptioned audio is not necessarily noise | ||
| or silence. A completed transcript does not establish that a recording contains | ||
| only microphone speech. Mono loudness uses FFmpeg's `dual_mono=true` playback | ||
| compensation consistently; it must not be mixed with uncompensated mono metrics. | ||
|
|
||
| ## Source and timing guarantees | ||
|
|
||
| The worker only reads an absolute regular local source, hashes it before and after, | ||
| and writes into a unique temporary directory. It copies video packets and verifies | ||
| them with the existing packet-proof helper. Existing finalization checks are | ||
| unchanged. Results carry source/output hashes, metrics, version, and validation | ||
| failures; every nonempty validation failure list disqualifies that candidate. | ||
|
|
||
| The worker restricts demuxers and protocols to local media files, rejecting playlists | ||
| instead of following their references. Only AAC inputs are eligible for processing; | ||
| other codecs are left unchanged. A MOV/PCM fixture exposed a video preservation | ||
| mismatch, so the first rollout deliberately bypasses that format. | ||
|
|
||
| The worker skips silence, extreme levels, existing clipping, unsupported formats, | ||
| already loud content, nonzero audio start times, discontinuous source timestamps, | ||
| and mismatched source audio/video durations. Voice processing additionally requires | ||
| `speechOnlyConfirmed`; the benchmark explicitly overrides this only for local | ||
| research. There is no production content classifier in this change. | ||
|
|
||
| FFmpeg's denoiser delays content by two sample-advance blocks without adjusting | ||
| PTS. Padding the tail and trimming that delay preserves boundary speech. Integer | ||
| sample timebases avoid timestamp rounding at 44.1 kHz. The encoded AAC result is | ||
| remeasured, with one bounded peak correction rendered from the original if needed. | ||
| Failed validation never authorizes publication. Cancellation, timeouts, and exceptions | ||
| clean up only the worker's own temporary files. | ||
|
|
||
| 46 tests cover policy gates, mono/stereo, 44.1/48 kHz, both profiles, speech-like | ||
| markers at clip boundaries, exact video packets, source preservation, silence, | ||
| nonzero/discontinuous timestamps, cancellation, and timeout. Scoped TypeScript and | ||
| Biome checks also pass. Measurements used macOS FFmpeg 8.0.1 and Bun 1.4.0; | ||
| the reviewed format suite also runs in both production-image architectures and | ||
| in the Railway Docker build. The full production cohort was rerun locally. | ||
|
|
||
| ## Reproducing | ||
|
|
||
| Keep source media, transcripts, per-recording measurements, and customer identifiers | ||
| outside the repository. Aggregate results and the frozen worker source hash are in | ||
| [audio-quality-benchmark-summary.json](audio-quality-benchmark-summary.json). | ||
| The local study retains `final-summary.json`, `final-voice-results.json`, | ||
| `final-levels-results.json`, and `report.md`, plus per-recording run receipts. | ||
| Interrupted runs and retries are retained separately. | ||
|
|
||
| The manifest is a JSON array with `id`, `split` (`tuning` or `holdout`), `stratum`, | ||
| `createdAt`, and `duration`. Sources are `sources/<id>.m4a`; transcripts are | ||
| `sources/<id>.vtt`. Initial download uses the authenticated Cap CLI for existing | ||
| transcripts and the public playlist for audio. No new transcription is requested. | ||
|
|
||
| ```sh | ||
| python3 scripts/benchmark-instant-audio.py /absolute/study --phase baseline | ||
| python3 scripts/benchmark-instant-audio.py /absolute/study --phase tuning --policies gain6 gain12 dynamic equalized clean | ||
| bun apps/media-server/scripts/benchmark-audio-quality.ts /absolute/study tuning unique-label voice | ||
| bun apps/media-server/scripts/benchmark-audio-quality.ts /absolute/study holdout another-label levels | ||
| ``` | ||
|
|
||
| Use a new label per run; the worker benchmark will not overwrite existing results. | ||
| Rejected candidates are recorded with their validation failures but are not copied | ||
| into the output set. Earlier historical runs retained rejected files for diagnosis. | ||
| The intelligibility calibration requires NumPy, SciPy, and pystoi. Supply three | ||
| reference IDs with `--reference-ids`; their M4A files must be two directories above | ||
| the output directory. Output-directory suffix `-v2` selects the corrected mild | ||
| policy; a name containing `strength` selects the 6/12 dB comparison. This calibration | ||
| script records historical filter alternatives; the TypeScript worker benchmark is | ||
| the authoritative final implementation. | ||
|
|
||
| ## Production-data revalidation | ||
|
|
||
| The tighter local-input restrictions were applied to the original 60-recording | ||
| cohort again. All 40 accepted outputs had identical hashes to the benchmark outputs; | ||
| 20 sources were left unchanged. Sixteen additional public production files, including | ||
| browser captures and recordings without completed transcripts, were left unchanged | ||
| by stream, level, headroom, or timestamp gates. All original hashes were preserved. | ||
| These are bounded compatibility checks, not proof of safety for every possible file. | ||
|
|
||
| ## Before serving any enhanced audio | ||
|
|
||
| Human review of the 12 volume-matched A/B excerpts is still required. Speech-only | ||
| eligibility, noisy and mixed-system-audio cases, and recordings excluded by the | ||
| public/transcribed selection need broader coverage. The short holdout clip rejected | ||
| for excessive gain must remain on its original audio; do not relax its gate to make | ||
| the benchmark pass. | ||
|
|
||
| Run the exact policy in the production Linux image, then verify actual share-page, | ||
| embed, seeking, downloads, edits, transcript alignment, and fallback behavior. | ||
| Measure worker memory, throughput, storage, and tail latency before rollout. | ||
|
|
||
| Future integration should create a separately versioned derivative after the | ||
| original is available, using a durable idempotent job bound to the source hash. | ||
| Publish atomically only after validation and only if the source still matches. | ||
| Keep the original available throughout processing, on failure, and for rollback. | ||
| Existing desktop installs could then benefit server-side without a capture update; | ||
| this experiment does not yet implement that serving integration. | ||
Uh oh!
There was an error while loading. Please reload this page.