Skip to content

Stop timestamps going backwards during decoding - #16

Merged
CodeWithBehnam merged 1 commit into
mainfrom
claude/lucid-franklin-gsu7v4
Sep 30, 2026
Merged

CodeWithBehnam merged 1 commit into
mainfrom
claude/lucid-franklin-gsu7v4

Conversation

@CodeWithBehnam

Copy link
Copy Markdown
Owner

What does this PR do?

Closes #6.

ApplyTimestampRules collected the positions of timestamp tokens in the sampled sequence rather than their values. The mask slice timestamp_begin:<position> was therefore always empty, and the rule never applied. The model could emit a timestamp earlier than the previous one, or close a segment at the timestamp it opened with.

This uses the token values, as OpenAI's reference implementation does. Every timestamp below the last one is masked. The last one is masked too, unless the last token is a lone timestamp closing a segment, so that <|2.00|><|2.00|> can still close one segment and open the next.

This changes decoded output: it is now closer to the reference behaviour.

Test helpers: random test models are cast to float16, as the published checkpoints are, and only floating-point parameters are cast (alignment_heads stays an integer array). The warning tests from #14 now use the plain stub model, which is faster.

How was this tested?

  • Tested with audio file(s)
  • Ran existing tests (pytest): 34 passed
  • Tested CLI (vayu audio.mp3)

tests/test_timestamp_rules.py:

  • Unit checks on the mask after text, after a lone closing timestamp, and after a pair.
  • A decode with a random-weight model, asserting timestamps never decrease and segments have non-zero length.

Three of these fail on main.

claude-review will fail as on #12 (the repository's CLAUDE_CODE_OAUTH_TOKEN secret, see this comment).

🤖 Generated with Claude Code

https://claude.ai/code/session_01TMXYMqgLAykqRApmbRfpTA


Generated by Claude Code

ApplyTimestampRules collected the positions of timestamp tokens in the
sampled sequence rather than their values, so the mask slice
timestamp_begin:<position> was always empty and the rule never applied:
the model could emit a timestamp earlier than the previous one, or close
a segment at the timestamp it opened with.

Use the token values, as OpenAI's reference implementation does: mask
every timestamp below the last one, and the last one itself unless the
last token is a lone timestamp closing a segment (so "<|2.00|><|2.00|>" can
still close one segment and open the next).

Tests: unit checks on the mask, and a decode with a random-weight model
asserting timestamps never decrease. Random test models are now cast to
float16 like the published checkpoints.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TMXYMqgLAykqRApmbRfpTA
@CodeWithBehnam
CodeWithBehnam merged commit f776cb5 into main Sep 30, 2026
3 of 4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Timestamp rule that stops timestamps going backwards never applies

2 participants