Document local models and batched-mode trade-offs; add a benchmark - #21
Merged
Merged
Conversation
- README: how to load models from local directories (WHISPER_MLX_MODEL_DIRS and the default allowed locations), which was undocumented and failed with no pointer to the fix. - README: what changes with batch_size > 1 (prompt conditioning per batch, fixed windows, hallucination_silence_threshold and best_of not applied). - scripts/benchmark.py times batch_size=1 against a chosen batch size on a given file and model, after loading and warming up the model, so the speed-up can be measured on the user's own hardware. The README points to it next to the batch size recommendations. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TMXYMqgLAykqRApmbRfpTA
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What does this PR do?
Closes #11.
/usr/local/share/whisper-mlx, or directories listed inWHISPER_MLX_MODEL_DIRS. This was undocumented, and the error gave no pointer to the fix. The allow-list itself is unchanged; whether to keep it is the maintainer's call (see Docs: model directory allow-list, batched-mode limits and speed claim #11).batch_size > 1:hallucination_silence_thresholdandbest_ofare not appliedscripts/benchmark.pytimesbatch_size=1against a chosen batch size on a given file and model. It loads and warms up the model first and decodes the audio once up front, so the timings cover transcription only. The README points to it next to the batch size recommendations, so the speed-up can be measured on the user's own Mac.How was this tested?
scripts/benchmark.pyend to end on 65 s of generated audio decoded by a real ffmpeg, using a random-weight model saved as a local model directory. WithoutWHISPER_MLX_MODEL_DIRSit fails with exactly the error the README quotes; with it, both decoding paths run. The timings were meaningless (CPU, random weights) and aren't reported.pytest): 80 passed, 1 skippedvayu audio.mp3)claude-reviewwill fail as on #12 (the repository'sCLAUDE_CODE_OAUTH_TOKENsecret, see this comment).🤖 Generated with Claude Code
https://claude.ai/code/session_01TMXYMqgLAykqRApmbRfpTA
Generated by Claude Code