Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 2 additions & 1 deletion docs/DESIGN-RULES.md
Original file line number Diff line number Diff line change
Expand Up @@ -115,7 +115,8 @@ an explicit cue that thinking time is expected. A prompt without a pause is a li
## 3. Sentence-level style (`SENT-*`, `ORI-*`, `SIG-*`, `VOI-*`)

**SENT-01** Median sentence length **≤ 20 words**; hard cap **35 words**; no sentence with more than
two subordinate clauses. Long source sentences are split, not compressed. *(KB §2.1; lint)*
two subordinate clauses. Long source sentences are split, not compressed. *(KB §2.1; lint; stage 6
`split`, which is checked in both directions because a split is lossless and a summary is not)*

**SENT-02** **No unresolved anaphora across a beat boundary.** "It", "this", "the former/latter",
"the above" must be replaced by the referent whenever the antecedent is more than one sentence back
Expand Down
4 changes: 2 additions & 2 deletions docs/PLAN-ai.md
Original file line number Diff line number Diff line change
Expand Up @@ -31,7 +31,7 @@ feature nobody can turn on.

| Where | What the model does | Provider | Ever run? |
|---|---|---|---|
| Stage 6, `elaborate` | gloss, anchor, analogy, why, compress | Anthropic only | **No** |
| Stage 6, `elaborate` | gloss, anchor, analogy, why, compress, split | Anthropic only | **No** |
| Stage 6, figures | describe a rendered crop | Anthropic only | **No** |
| Stage 6, gate | entailment check (`GRD-02`) | Anthropic only | **No** |
| MCP assistant | *all of the above*, written in the conversation | none needed | **Yes** |
Expand Down Expand Up @@ -60,7 +60,7 @@ should stop treating them as one thing.

**Author.** The model produces text that reaches the listener. Highest risk — a fluent wrong
sentence is the worst output this system can make — and every output must pass the grounding
gate. Currently: gloss, anchor, analogy, why, compress, figure.
gate. Currently: gloss, anchor, analogy, why, compress, figure, split.

**Critic.** The model reads output the deterministic pipeline produced and says what is wrong
with it. Low risk: its output is a report, not a script, and a wrong criticism costs a human
Expand Down
20 changes: 19 additions & 1 deletion docs/ai/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,7 +8,8 @@ icon: lucide/brain-circuit
fluency, and the structure is deterministic: a document goes in and a listenable, checkable
programme comes out with no model involved at all. What a model adds is the explaining — a
one-line gloss, a concrete anchor, an analogy that says where it breaks down, a description of a
figure you cannot see.
figure you cannot see — and one thing that is not explaining at all: cutting the paper's longest
sentences into ones you can hold in your head.

Start here:

Expand Down Expand Up @@ -71,6 +72,23 @@ sentence must appear in the source sentences it was written from; a claim that r
paper's direction is rejected before it reaches the programme (`GRD-03`). You can watch that
happen — the elaboration report lists what was accepted, what was rejected and why.

**One task rewrites the paper's own words, and only one.** Rule `SENT-01` caps a sentence at
thirty-five spoken words, because a sentence you would re-read on the page is simply lost in
audio — and 465 sentences across a twelve-paper corpus are over it. The rule says what to do
about them: *long source sentences are split, not compressed.*

That distinction is the whole of it. A summary of a sentence reads exactly like a split of it,
and the listener has no way to tell which they were given. So the check runs in **both**
directions: every number in the source must appear in the split, and every number in the split
must appear in the source. Nothing else a model writes here can fail the first of those, because
nothing else is supposed to be lossless. It is what makes this safe to do to a paper's sentences
at all.

The paper is never edited. The split is stored beside the sentence it replaces, the span still
points at what was written, and `study.md` can show you one against the other. When a split
drops a value it is rejected and the long sentence is spoken as it stands — which happened twice
in ten on the first live run, and a long true sentence beats a short lossy one every time.

**When the model is absent, every task degrades along a documented path** and the manifest
records *why* — there are four different reasons and they need four different actions from you:
no provider configured, configured but unreachable, reached but refused the shape, or out of
Expand Down
11 changes: 10 additions & 1 deletion src/mimem/config.py
Original file line number Diff line number Diff line change
Expand Up @@ -104,6 +104,15 @@ class ElaborationBudget(BaseModel):
max_anchors: int = 4
anchor_min_abstractness: float = 0.50

#: Sentences to split in one build (rule SENT-01), longest first.
#:
#: A cap on *calls*, unlike everything above it, and it has to be: the others scale with the
#: number of concepts a paper has and this one scales with the paper. A ninety-eight page
#: review offers well over a thousand sentences past the cap, and splitting all of them would
#: cost more than every other task in this file put together while the four analogies that
#: carry the programme went unwritten.
max_splits: int = 40

#: Which implementation answers each task: ``off``, ``assist`` or ``prefer``. See
#: :mod:`mimem.elaborate.reconcile`.
#:
Expand All @@ -121,7 +130,7 @@ class ElaborationBudget(BaseModel):
#: keep a paid run away from it entirely.
modes: dict[str, str] = Field(
default_factory=lambda: dict.fromkeys(
("gloss", "anchor", "analogy", "why", "figure", "compress"), "prefer"
("gloss", "anchor", "analogy", "why", "figure", "compress", "split"), "prefer"
)
)

Expand Down
108 changes: 97 additions & 11 deletions src/mimem/elaborate/run.py
Original file line number Diff line number Diff line change
Expand Up @@ -31,9 +31,11 @@

from mimem.config import Listener, Profile
from mimem.elaborate.reconcile import Absence, Deterministic, Mode, classify
from mimem.elaborate.sentences import LongSentence, over_long, verify_split
from mimem.ir import (
Analogy,
Anchor,
Block,
Concept,
ConceptRegistry,
Document,
Expand All @@ -44,7 +46,7 @@
from mimem.llm.cache import CacheStats
from mimem.llm.client import Client, LLMRefusedError, LLMUnavailableError, NullClient, Request
from mimem.llm.cost import BudgetExceededError, Ledger, Plan, estimate
from mimem.llm.schemas import AnalogyOut, AnchorOut, FigureOut, GlossOut, WhyOut
from mimem.llm.schemas import AnalogyOut, AnchorOut, FigureOut, GlossOut, SplitOut, WhyOut

# The supporting-sentence index is shared between stage 6 and stage 7. It lives with the
# planner, which is its heavier user; importing it here is deliberate rather than a layering
Expand All @@ -57,8 +59,14 @@
#: need to: the task carries the sentences that matter, and the document is context.
MAX_DOCUMENT_CHARS = 60_000

#: The cap the split is asked to get under, and it is *below* the linter's thirty-five. A
#: sentence that lands exactly on the limit written is over it spoken, because the numbers in it
#: have not been said yet -- see :func:`~mimem.elaborate.sentences.over_long`. Asking for
#: twenty-five leaves room for the verbalizer.
SPLIT_CAP = 25

#: Every task that can be reconciled, so a mode exists for each.
TASKS = ("gloss", "anchor", "analogy", "why", "figure", "compress")
TASKS = ("gloss", "anchor", "analogy", "why", "figure", "compress", "split")


@dataclass(frozen=True)
Expand Down Expand Up @@ -217,6 +225,12 @@ def _requests(
for task in figure_tasks(doc):
yield tasks.figure(task.caption, task.references, document)

# Sentence splits, which are the one task whose count scales with the length of the paper
# rather than with the number of concepts -- so a dry run that did not price them was
# quoting for the wrong build on anything longer than a letter.
for sentence in over_long(doc, profile, listener)[: profile.elaboration.max_splits]:
yield tasks.split(sentence.text, document, SPLIT_CAP)


def elaborate(
doc: Document,
Expand Down Expand Up @@ -275,7 +289,7 @@ def elaborate(
report,
client,
tasks.gloss(concept, support, document, listener),
concept,
concept.canonical,
partial(_apply_gloss, concept),
source=source,
fallback="the source's own definitional sentence",
Expand All @@ -289,7 +303,7 @@ def elaborate(
report,
client,
tasks.anchor(concept, support, document, listener),
concept,
concept.canonical,
partial(_apply_anchor, concept),
source=source,
kinds=GROUNDING_KINDS["anchor"],
Expand All @@ -303,7 +317,7 @@ def elaborate(
report,
client,
tasks.why(concept, support, document),
concept,
concept.canonical,
partial(_apply_why, concept, spans),
source=source,
fallback="no why-explanation",
Expand All @@ -320,7 +334,7 @@ def elaborate(
report,
client,
tasks.analogy(concept, support, document, listener),
concept,
concept.canonical,
partial(_apply_analogy, concept),
source="\n".join(support),
kinds=GROUNDING_KINDS["analogy"],
Expand All @@ -331,9 +345,69 @@ def elaborate(
)

_describe_figures(report, client, doc, document, out_dir, model=model, progress=progress)
_split_sentences(
report,
client,
doc,
document,
profile,
listener,
mode=modes["split"],
model=model,
progress=progress,
)
return report


def _split_sentences(
report: ElaborationReport,
client: Client,
doc: Document,
document: str,
profile: Profile,
listener: Listener | None,
*,
mode: Mode,
model: str | None,
progress: Callable[[str, str], None] | None,
) -> None:
"""Cut the sentences a listener cannot hold in one piece (rule SENT-01).

Last, and deliberately. Every other task competes for the elaboration budget against the
*concepts* rule DIF-02 ranks; this one competes against the length of the paper, and a
hundred splits would starve the four analogies that carry the programme. Running it after
the others means the budget answers the question in the right order.
"""
if mode is Mode.OFF:
return
for sentence in over_long(doc, profile, listener)[: profile.elaboration.max_splits]:
block = doc.block(sentence.block_id)
_run(
report,
client,
tasks.split(sentence.text, document, SPLIT_CAP),
_shorten(sentence.text),
partial(_apply_split, block, sentence),
source=sentence.text,
inspect=lambda data, source: verify_split(source, list(data.sentences), cap=SPLIT_CAP),
fallback="the source sentence, unchanged",
model=model,
progress=progress,
mode=mode,
)


def _shorten(text: str, words: int = 6) -> str:
"""A sentence named by its opening, for the progress line and the degradation report."""
head = text.split()[:words]
return " ".join(head) + ("..." if len(text.split()) > words else "")


def _apply_split(block: Block, sentence: LongSentence, out: SplitOut) -> None:
"""Store the split beside the sentence it replaces, leaving the source alone."""
block.rewrites[sentence.key] = " ".join(s.strip() for s in out.sentences)


def _describe_figures(
report: ElaborationReport,
client: Client,
Expand Down Expand Up @@ -401,12 +475,13 @@ def _run(
report: ElaborationReport,
client: Client,
request: Request,
concept: Concept,
subject: str,
apply: Callable[[Any], None],
*,
source: str,
fallback: str,
kinds: tuple[str, ...] = ("number", "year", "name", "direction"),
inspect: Callable[[Any, str], list[Finding]] | None = None,
on_degrade: Deterministic | None = None,
mode: Mode = Mode.PREFER,
model: str | None = None,
Expand All @@ -418,6 +493,13 @@ def _run(
failure -- and under ``assist`` it runs *first*, which is the whole point of the mode: the
rules answer what they are good at, and the model is asked only about the rest. It returns
whether it produced anything, because nothing else can tell the caller that.

``inspect`` replaces the grounding check for a task whose contract is different. Everything
here writes *new* text and can only be asked whether it invented something; a sentence split
is lossless, so it is also asked whether it lost something, which no other task can fail.

``subject`` is a name for the thing being worked on, for the progress line and the
degradation report. It was a whole :class:`Concept` until the split arrived, which has none.
"""
deterministic = on_degrade

Expand All @@ -436,7 +518,7 @@ def _run(
if model:
request = replace(request, model=model)
if progress is not None:
progress(request.task, concept.canonical)
progress(request.task, subject)

report._count(report.attempted, request.task)
try:
Expand All @@ -445,19 +527,23 @@ def _run(
except BudgetExceededError:
raise
except (LLMUnavailableError, LLMRefusedError) as exc:
_degrade(report, request.task, concept.canonical, str(exc), fallback, classify(exc))
_degrade(report, request.task, subject, str(exc), fallback, classify(exc))
if deterministic is not None and deterministic():
report._count(report.deterministic, request.task)
return

report.ledger.record(response)
findings = ground(response.data, source, kinds=kinds)
findings = (
inspect(response.data, source)
if inspect is not None
else ground(response.data, source, kinds=kinds)
)
if findings:
report.rejected.extend(findings)
_degrade(
report,
request.task,
concept.canonical,
subject,
"; ".join(str(f) for f in findings[:3]),
fallback,
Absence.REJECTED,
Expand Down
Loading
Loading