Skip to content

Split the sentences a listener cannot hold, and check both directions - #47

Merged
jepegit merged 1 commit into
mainfrom
sentence-splits
Sep 9, 2026
Merged

jepegit merged 1 commit into
mainfrom
sentence-splits

Conversation

@jepegit

@jepegit jepegit commented Sep 9, 2026

Copy link
Copy Markdown
Owner

SENT-01 caps a sentence at 35 spoken words. 465 of the stress corpus's sentences are over it — the largest family of findings in the project, and the one the rule's own docstring says belongs to a rewriting stage.

The design rule says what that stage must do:

Long source sentences are split, not compressed.

Why a split is the one rewrite that can be checked

A summary of a sentence reads exactly like a split of it. The prose cannot tell you which you were given, so the check has to — and a split is lossless, which no other task here is:

direction catches can any other stage-6 task fail it?
every number in the source appears in the split a summary, a dropped condition no
every number in the split appears in the source an invented value yes
directions and names unchanged (GRD-03) "higher" for "lower" yes

That asymmetry is the entire basis for doing this to a paper's own words. A gloss and an anchor are new text and can only be asked whether they invented something; only a split can be asked whether it lost something.

Plus four cheaper checks: at least two sentences back, each under the cap, and no more than 1.5× growth — because a repeated subject costs words but half as much again is an expansion, and expansion is where facts get added.

The paper is never edited

The split is stored on Block.rewrites, keyed by the offsets of the sentence it replaces. The span still points at what the paper wrote — which is what GRD-02 checks the beat against, and what lets study.md show one against the other. Only what is spoken changes.

Two details that are easy to get wrong

Selection is measured on the verbalized text. "298 K" is two words written and four heard. Selecting on the written form misses the sentences that are long because of what they state, which are the ones most in need of splitting. The split is asked to get under 25 rather than 35 for the same reason: a sentence cut to exactly the cap in written words is over it once its numbers have been said.

The subject is repeated, not pronominalised. "X, which does Y" splits naturally into "X. It does Y" — a SENT-02 violation manufactured by the fix for SENT-01. The instruction says so, and the live output does it:

"Silicon has the highest theoretical specific energy density (4008 mAh/g for Li21Si5)… Silicon has yet to find use in commercial lithium-ion cells."

Capped at 40 sentences a build, and run last. Every other task scales with the number of concepts; this one scales with the paper, and a 98-page review offers over a thousand candidates. Running it last means the budget starves splits rather than the four analogies that carry the programme.

First live run

Haiku, one corpus paper, $0.09 for the whole stage (42 calls — the dry-run estimate was $0.09):

attempted: {'gloss': 12, 'why': 12, 'anchor': 4, 'analogy': 3, 'figure': 8, 'split': 10}
succeeded: {'gloss': 10, 'why': 7,  'anchor': 4,               'figure': 2, 'split': 8}
rejected : split — number '2.3': dropped by the split; a split may not lose a value
           split — number '0':   dropped by the split; a split may not lose a value

SENT-01 on that paper: 8 → 2, total warnings 26 → 12.

The gate caught two real failures on the first ten tries, and the sentences they came from are spoken as they stand. One of the two remaining SENT-01 findings is that rejected split — the system preferring a long true sentence to a short lossy one, which is the right preference. The other is a beat the model wrote itself; this stage splits the paper's sentences, not its own, and that gap is untouched here.

The deterministic path is unchanged

--local is still the default, over_long is pure, and the corpus is bit-identical without a model: 2 errors, 10/12 clean.

18 new tests in tests/unit/test_sentence_splits.py, scripted rather than recorded — they check the plumbing and the gate, neither of which should depend on what a model says today.

🤖 Generated with Claude Code

Rule SENT-01 caps a sentence at thirty-five spoken words and 465 of the stress
corpus's sentences are over it -- the largest family of findings in the project,
and the one the rule's own docstring says belongs to a rewriting stage. The design
rule says what that stage must do: long source sentences are split, not compressed.

That distinction is the whole design. A summary of a sentence reads exactly like a
split of it, and the listener has no way to know which they were given. So the
check runs in both directions: every number in the source must appear in the
split, and every number in the split must appear in the source. Nothing else stage
6 writes can fail the first of those, because nothing else is meant to be
lossless -- a gloss and an anchor are new text and can only be asked whether they
invented something. The asymmetry is the entire basis for doing this to a paper's
own words, and it is what the tests are about.

The paper is never edited. The split is stored beside the sentence it replaces,
keyed by the offsets it covers, and the span still points at what was written --
which is what rule GRD-02 checks the beat against and what lets study.md show one
against the other. Only what is spoken changes.

Selection is measured on the verbalized text, because "298 K" is two words written
and four heard, and the linter and the listener both get the second one. The split
is asked to get under twenty-five rather than thirty-five for the same reason:
a sentence cut to exactly the cap in written words is over it once its numbers
have been said.

The instruction asks for the subject to be repeated rather than pronominalised.
"X, which does Y" splits naturally into "X. It does Y", and that is a SENT-02
violation manufactured by the fix for SENT-01 -- the listener meets "it" with the
referent now behind a full stop.

Capped at forty sentences a build. Every other task here scales with the number of
concepts a paper has; this one scales with the paper, and a ninety-eight page
review offers over a thousand candidates. Running it last means the budget answers
in the right order rather than starving the four analogies that carry the
programme.

_run now takes a subject rather than a Concept, and an optional check of its own.
The split has no concept, and its contract is not the grounding contract.

First live run, Haiku, one corpus paper, $0.09 for the whole stage: ten splits
attempted, eight accepted, two rejected for dropping a number -- and the sentences
they came from are spoken as they stand. SENT-01 on that paper falls from eight to
two. One of the two is the rejected split, which is the gate preferring a long true
sentence to a short lossy one, and the other is a beat the model wrote itself,
which this stage does not cover.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@jepegit
jepegit merged commit cb365c1 into main Sep 9, 2026
4 checks passed
@jepegit
jepegit deleted the sentence-splits branch September 9, 2026 07:59
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant