Conversation
Every narrated render failed validation with audio around 3.3 of 5 seconds, and the cause was not the padding. -frames:v ends the whole output the moment the video stream reaches its limit, which cuts the audio wherever the encoder happened to have flushed to. That is why none of the obvious fixes worked. apad, apad=whole_dur=5, atrim=0:5 with asetpts, -shortest: all of them left the audio short, and an input already padded to exactly 5.000s still came out at 3.264s. The padding was fine the whole time; the output was being terminated early. An encode that carries sound is now bounded by duration instead. The frame count is still exactly 150, because 150 frames at 30fps is five seconds — the guarantee now comes from arithmetic rather than from a flag, and validation counts the frames either way. Silent encodes keep -frames:v, since they have no audio to truncate and a frame count cannot be satisfied by a timebase rounding error the way a duration can. -shortest goes from the music path as well: with both streams bounded to the same explicit length there is no shorter one to find. Worth recording why this took so long. The behaviour is version dependent: on ffmpeg 8, which is what this repo's test suite runs against locally, the narrated encode produces a correct 5.000s track and the test I added passes. On ffmpeg 4.4.2, which is what ships in the image, the same arguments produce 3.306s. A test that runs only on the newer binary cannot catch this, so it was reproduced by running ffmpeg inside the deployed container. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
ThreatCrush Security Scan48 finding(s) HIGH/CRITICAL: 2 | MEDIUM: 31 | LOW: 15
Snippets are redacted; ThreatCrush never prints matched credential material. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Every narrated render failed validation with audio around 3.3 of 5 seconds. The cause was not the padding.
-frames:vends the whole output the moment the video stream reaches its limit, which cuts the audio wherever the encoder happened to have flushed to.Why none of the obvious fixes worked
Reproduced inside the deployed container, one at a time:
-af apad -shortest-af apad -t 5-af "apad,atrim=0:5,asetpts=N/SR/TB"-af apad=whole_dur=5-frames:v,-af apad -t 5The fifth row is the one that settles it: an input already padded to exactly five seconds still came out short. The padding was fine the whole time — the output was being terminated early.
The change
An encode that carries sound is bounded by duration instead. The frame count is still exactly 150, because 150 frames at 30fps is five seconds — the guarantee now comes from arithmetic rather than a flag, and validation counts the frames either way (the test asserts 150 decoded frames).
Silent encodes keep
-frames:v: they have no audio to truncate, and a frame count cannot be satisfied by a timebase rounding error the way a duration can. That distinction is now tested in both directions.-shortestgoes from the music-bed path too — with both streams bounded to the same explicit length, there is no shorter one to find. I updated #293's assertion rather than deleting it, with the reasoning.Why this took several attempts, honestly
The behaviour is version dependent. On ffmpeg 8 — what this repo's suite runs against locally — the narrated encode produces a correct 5.000s track, and the test I added for it passes. On ffmpeg 4.4.2, which is what ships in the image, the same arguments produce 3.306s.
So a green local suite was not evidence, and I treated it as evidence twice. This was finally found by running ffmpeg inside the deployed container and bisecting the arguments.
Verification
2590 passed / 1 failed repo-wide — the pre-existing
tracker-geofailure. Typechecks clean. The decisive check is on dev2 after deploy, against 4.4.2, and I'll report the probed result rather than the suite's.