Skip to content

Keep text after the last timestamp pair in batched decoding - #13

Merged
CodeWithBehnam merged 1 commit into
mainfrom
claude/lucid-franklin-gsu7v4
Sep 30, 2026
Merged

CodeWithBehnam merged 1 commit into
mainfrom
claude/lucid-franklin-gsu7v4

Conversation

@CodeWithBehnam

Copy link
Copy Markdown
Owner

What does this PR do?

Closes #3.

  • With batch_size > 1, text decoded after the last pair of timestamp tokens in a window was dropped. The one-window path re-decodes that audio from the last timestamp, but a batch always moves on by a full window, so the text was never recovered. It's now kept as a final segment from its timestamp to the end of the window. The last window ends at the end of the audio.
  • Each segment from a batch recorded the batch's first seek. new_segment now takes seek explicitly: the window's own offset in the batched path, and unchanged in the sequential path.
  • Adds tests/conftest.py with a stub model, so transcribe() can be tested without weights.

How was this tested?

  • Tested with audio file(s)
  • Ran existing tests (pytest): 19 passed
  • Tested CLI (vayu audio.mp3)

New tests in tests/test_transcribe_batched.py. Two fail on main: "world" is dropped, and seek values come out as [0, 0, 0]. Three pin behaviour that shouldn't change: a single-timestamp ending, a trailing timestamp with no text, and text with no timestamp pairs.

The claude-review check will fail as it did on #12: the CLAUDE_CODE_OAUTH_TOKEN secret doesn't authenticate (details in this comment).

🤖 Generated with Claude Code

https://claude.ai/code/session_01TMXYMqgLAykqRApmbRfpTA


Generated by Claude Code

With batch_size > 1, text decoded after the last pair of timestamp tokens
in a window was dropped: the one-window path re-decodes that audio from
the last timestamp, but a batch always moves on by a full window, so the
text was never recovered. It is now kept as a final segment that runs to
the end of its window.

Each segment from a batch also recorded the batch's first seek instead of
its own window's; new_segment now takes seek explicitly.

Adds a stub-model fixture so transcribe() can be exercised without weights.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TMXYMqgLAykqRApmbRfpTA
@CodeWithBehnam
CodeWithBehnam merged commit 8e5df31 into main Sep 30, 2026
3 of 4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Batched decoding drops text at the end of each 30 s window

2 participants