Merge a segment too short to be one - #46
Merged
Merged
Conversation
Rule SEG-01 has two ends and the planner only enforced the top one. Five of the stress corpus's segments ran under sixteen seconds and one ran seven -- a position statement and a transition, with no teaching beat in it at all. A listener does not experience that as a segment. They experience two transitions, seven seconds apart, with a sentence between them. The mirror of _split_oversized, and it rests on the same argument that function already makes: the closing recap, the emphasis marker and the spaced callbacks all land after segmentation, so a segment's finished length is only known here. Three things it has to respect. The last segment of a section is left alone. It carries the recap and the section-boundary pause, and SEG-01 exempts it for that reason. The head's transition goes with the boundary it announced -- and it is not the head's last beat, though _segment put it there, because the segment prompt and its answer were appended after it at step 8. And the new-term budget is kept rather than traded. The first version of this pushed one corpus segment past SEG-03's limit, which that rule calls an error because a planner that packs a segment could have split it; turning a pacing warning into a failed build is not a trade worth making. The test is whether the merge makes anything worse, not whether the result is inside the budget: folding seven seconds of scaffolding that teaches nothing into a segment already over the budget adds nothing to it, and refusing on the absolute count left exactly those where they were. Corpus: segments under the floor fall from twenty to twelve, and every one under thirty-one seconds is gone. Errors unchanged at two, with ten of twelve clean. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
SEG-01has two ends and the planner only enforced the top one. Five of the corpus's segments ran under sixteen seconds; one ran seven — a position statement and a transition, with no teaching beat in it at all.A listener does not experience that as a segment. They experience two transitions, seven seconds apart, with a sentence between them.
This is the mirror of
_split_oversized, and it rests on the argument that function already makes: the closing recap, the emphasis marker and the spaced callbacks all land after segmentation, so a segment's finished length is only known here.Three things it has to respect
The last segment of a section is left alone. It carries the recap and the section-boundary pause, and
SEG-01exempts it for that reason.The head's transition goes with the boundary it announced — and it is not the head's last beat, though
_segmentput it there, because the segment prompt and its answer were appended after it at step 8. The test for this asserts one transition per segment, not "a transition last", since a transition third-from-last is the ordinary shape.The new-term budget is kept, not traded. The first version pushed one corpus segment past
SEG-03's limit — which that rule calls an error, because a planner that packs a segment could have split it. Turning a pacing warning into a failed build is not a trade worth making.But the test is whether the merge makes anything worse, not whether the result is inside the budget. The commonest short segment teaches nothing at all, and folding it into a segment already over the budget adds nothing to it — refusing on the absolute count left exactly those where they were. That cost a corpus round-trip to notice.
Corpus
Every segment under 31 seconds is gone. The twelve that remain are 31–45 s and cannot merge without breaking the 120 s ceiling or the term budget.
The
SEG-03mutator intest_script_lint.pynow chooses its segment instead of takingsegments[0]: a segment already holding a beat with more new terms than the budget reports as a warning whatever you add to it, so mutating that one tests nothing — andsegments[0]became such a segment the day this merge landed. The assertion message says so.🤖 Generated with Claude Code