Split the sentences a listener cannot hold, and check both directions - #47
Merged
Merged
Conversation
Rule SENT-01 caps a sentence at thirty-five spoken words and 465 of the stress corpus's sentences are over it -- the largest family of findings in the project, and the one the rule's own docstring says belongs to a rewriting stage. The design rule says what that stage must do: long source sentences are split, not compressed. That distinction is the whole design. A summary of a sentence reads exactly like a split of it, and the listener has no way to know which they were given. So the check runs in both directions: every number in the source must appear in the split, and every number in the split must appear in the source. Nothing else stage 6 writes can fail the first of those, because nothing else is meant to be lossless -- a gloss and an anchor are new text and can only be asked whether they invented something. The asymmetry is the entire basis for doing this to a paper's own words, and it is what the tests are about. The paper is never edited. The split is stored beside the sentence it replaces, keyed by the offsets it covers, and the span still points at what was written -- which is what rule GRD-02 checks the beat against and what lets study.md show one against the other. Only what is spoken changes. Selection is measured on the verbalized text, because "298 K" is two words written and four heard, and the linter and the listener both get the second one. The split is asked to get under twenty-five rather than thirty-five for the same reason: a sentence cut to exactly the cap in written words is over it once its numbers have been said. The instruction asks for the subject to be repeated rather than pronominalised. "X, which does Y" splits naturally into "X. It does Y", and that is a SENT-02 violation manufactured by the fix for SENT-01 -- the listener meets "it" with the referent now behind a full stop. Capped at forty sentences a build. Every other task here scales with the number of concepts a paper has; this one scales with the paper, and a ninety-eight page review offers over a thousand candidates. Running it last means the budget answers in the right order rather than starving the four analogies that carry the programme. _run now takes a subject rather than a Concept, and an optional check of its own. The split has no concept, and its contract is not the grounding contract. First live run, Haiku, one corpus paper, $0.09 for the whole stage: ten splits attempted, eight accepted, two rejected for dropping a number -- and the sentences they came from are spoken as they stand. SENT-01 on that paper falls from eight to two. One of the two is the rejected split, which is the gate preferring a long true sentence to a short lossy one, and the other is a beat the model wrote itself, which this stage does not cover. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
SENT-01caps a sentence at 35 spoken words. 465 of the stress corpus's sentences are over it — the largest family of findings in the project, and the one the rule's own docstring says belongs to a rewriting stage.The design rule says what that stage must do:
Why a split is the one rewrite that can be checked
A summary of a sentence reads exactly like a split of it. The prose cannot tell you which you were given, so the check has to — and a split is lossless, which no other task here is:
GRD-03)That asymmetry is the entire basis for doing this to a paper's own words. A gloss and an anchor are new text and can only be asked whether they invented something; only a split can be asked whether it lost something.
Plus four cheaper checks: at least two sentences back, each under the cap, and no more than 1.5× growth — because a repeated subject costs words but half as much again is an expansion, and expansion is where facts get added.
The paper is never edited
The split is stored on
Block.rewrites, keyed by the offsets of the sentence it replaces. The span still points at what the paper wrote — which is whatGRD-02checks the beat against, and what letsstudy.mdshow one against the other. Only what is spoken changes.Two details that are easy to get wrong
Selection is measured on the verbalized text. "298 K" is two words written and four heard. Selecting on the written form misses the sentences that are long because of what they state, which are the ones most in need of splitting. The split is asked to get under 25 rather than 35 for the same reason: a sentence cut to exactly the cap in written words is over it once its numbers have been said.
The subject is repeated, not pronominalised. "X, which does Y" splits naturally into "X. It does Y" — a
SENT-02violation manufactured by the fix forSENT-01. The instruction says so, and the live output does it:Capped at 40 sentences a build, and run last. Every other task scales with the number of concepts; this one scales with the paper, and a 98-page review offers over a thousand candidates. Running it last means the budget starves splits rather than the four analogies that carry the programme.
First live run
Haiku, one corpus paper, $0.09 for the whole stage (42 calls — the dry-run estimate was $0.09):
SENT-01on that paper: 8 → 2, total warnings 26 → 12.The gate caught two real failures on the first ten tries, and the sentences they came from are spoken as they stand. One of the two remaining
SENT-01findings is that rejected split — the system preferring a long true sentence to a short lossy one, which is the right preference. The other is a beat the model wrote itself; this stage splits the paper's sentences, not its own, and that gap is untouched here.The deterministic path is unchanged
--localis still the default,over_longis pure, and the corpus is bit-identical without a model: 2 errors, 10/12 clean.18 new tests in
tests/unit/test_sentence_splits.py, scripted rather than recorded — they check the plumbing and the gate, neither of which should depend on what a model says today.🤖 Generated with Claude Code