Skip to content

Document local models and batched-mode trade-offs; add a benchmark - #21

Merged
CodeWithBehnam merged 1 commit into
mainfrom
claude/lucid-franklin-gsu7v4
Sep 30, 2026
Merged

CodeWithBehnam merged 1 commit into
mainfrom
claude/lucid-franklin-gsu7v4

Conversation

@CodeWithBehnam

Copy link
Copy Markdown
Owner

What does this PR do?

Closes #11.

  • README: Loading Local Models. Local model directories are only loaded from the HuggingFace cache, /usr/local/share/whisper-mlx, or directories listed in WHISPER_MLX_MODEL_DIRS. This was undocumented, and the error gave no pointer to the fix. The allow-list itself is unchanged; whether to keep it is the maintainer's call (see Docs: model directory allow-list, batched-mode limits and speed claim #11).
  • README: Batched vs Sequential Decoding. What changes with batch_size > 1:
    • prompt conditioning is per batch, not per window
    • windows are fixed, so text at a boundary becomes a segment ending there
    • hallucination_silence_threshold and best_of are not applied
    • fallback re-decodes only the windows that fail
  • scripts/benchmark.py times batch_size=1 against a chosen batch size on a given file and model. It loads and warms up the model first and decodes the audio once up front, so the timings cover transcription only. The README points to it next to the batch size recommendations, so the speed-up can be measured on the user's own Mac.

How was this tested?

  • Tested with audio file(s): ran scripts/benchmark.py end to end on 65 s of generated audio decoded by a real ffmpeg, using a random-weight model saved as a local model directory. Without WHISPER_MLX_MODEL_DIRS it fails with exactly the error the README quotes; with it, both decoding paths run. The timings were meaningless (CPU, random weights) and aren't reported.
  • Ran existing tests (pytest): 80 passed, 1 skipped
  • Tested CLI (vayu audio.mp3)

claude-review will fail as on #12 (the repository's CLAUDE_CODE_OAUTH_TOKEN secret, see this comment).

🤖 Generated with Claude Code

https://claude.ai/code/session_01TMXYMqgLAykqRApmbRfpTA


Generated by Claude Code

- README: how to load models from local directories
  (WHISPER_MLX_MODEL_DIRS and the default allowed locations), which was
  undocumented and failed with no pointer to the fix.
- README: what changes with batch_size > 1 (prompt conditioning per
  batch, fixed windows, hallucination_silence_threshold and best_of not
  applied).
- scripts/benchmark.py times batch_size=1 against a chosen batch size on
  a given file and model, after loading and warming up the model, so the
  speed-up can be measured on the user's own hardware. The README points
  to it next to the batch size recommendations.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TMXYMqgLAykqRApmbRfpTA
@CodeWithBehnam
CodeWithBehnam merged commit db3c74c into main Sep 30, 2026
3 of 4 checks passed
@CodeWithBehnam CodeWithBehnam mentioned this pull request Oct 1, 2026
1 of 3 tasks
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Docs: model directory allow-list, batched-mode limits and speed claim

2 participants