diff --git a/docs/expression-contract.md b/docs/expression-contract.md index 7e62306..495ae0a 100644 --- a/docs/expression-contract.md +++ b/docs/expression-contract.md @@ -21,13 +21,17 @@ route were not — and `decision-framing.md` mandated the exact opposite placement rule with no scope note, so two contradictory disclosure policies coexisted silently. -Three registries live here: +§3 is the **mother chapter**: the one shape every answer takes. It is not a +registry entry and has no ID, because it is the law the registries are +parameters of. Read it before anything below it. + +Three registries live under it: | Series | Owns | Where the rules are | |---|---|---| | **V1–V9** | Voice: what an answer leads with, what it may not manufacture, when it stops. | [output-voice.md](output-voice.md) — routed, never copied. | -| **D1–D7** | Disclosure relevance and placement: what materially qualifies a claim, and which floor it lives on. | §3 below. | -| **C1–C4** | Citation and provenance: how a claim says where it came from. | §4 below. | +| **D1–D7** | Disclosure relevance and placement: what materially qualifies a claim, and which floor it lives on. | §4 below. | +| **C1–C4** | Citation and provenance: how a claim says where it came from. | §5 below. | Voice stays in its own file rather than being folded in here because V1–V9 are a stable, ID-addressed registry with their own oracle @@ -64,7 +68,169 @@ Out of scope here, and unchanged by this file: cheaply without limiting relevant research, tools, or presentation. Nothing here licenses dropping a fact to satisfy a style preference; see D5. -## 3. Disclosure relevance and placement (D1–D7) +## 3. The answer pyramid — the mother law + +Owner ruling, 2026-08-21 +([#832](https://github.com/atomchung/fomo-kernel/issues/832)). **Every +user-visible answer this product gives has one shape.** It is mandatory, it is +small, and every surface below derives from it rather than restating it. D1–D7, +C1–C4, and every surface document are parameters of this shape. + +**Why one shape, and why now.** #830 deleted the obligation whitelist. The +post-merge rerun of the same four frozen scenes (receipts on +[PR #831](https://github.com/atomchung/fomo-kernel/pull/831)) shows composition +fixed — no basis recital, no rendered machine anchor, disclosures collapsed, +answers opening on stance — and **first-answer length inside ±5% of the +pre-trim arm in all four scenes**. Deleting obligations was necessary and not +sufficient: nothing positive said what an answer *is*, so the space the +deleted obligations vacated refilled with discretionary elaboration. Worse, +the one norm that could have said it already existed **five times, written +five different ways**: the card's keynote (output-contract.md §2), V1 +(output-voice.md), the reader-order paragraph (SKILL.md), "lead with the +bounded value" (decision-framing.md), "the host shows value first" +(weekly-market-read.md). Five local phrasings of one idea drift by +construction. #823 unified where the *rules* live and classified answer shape +as per-surface layout; that classification is the seam this chapter closes. + +Four external traditions converge on the same answer, and each contributed one +piece: the **consulting pyramid** (answer on top, supports under it, evidence +at the bottom, and the reader descends only as far as they need); the +**sell-side research note** (the first page is the rating, the thesis and the +numbers that carry it, and the anatomy is a house template rather than an +author's preference); **Anthropic's own prompting guidance** (worked examples +steer format more reliably than a list of constraints, and a small mandatory +base plus opt-in specialization beats one long rule set); **design-system token +layering** (a mandatory core, an additive component layer, a fall-through +default, and a different owner and cadence per layer). + +### 3.1 The four floors + +| Floor | What it carries | How much of it | +|---|---|---| +| **Top** | The answer: the stance, and the reason that decides it. | One sentence. Always present. | +| **Middle** | Only support that could change the decision. | As many blocks as pass the increment gate, and no more. | +| **Bottom** | Evidence, alternatives, and the rest of the computed inventory. | Not in the answer. It lives in the data layer and surfaces on request. | +| **End** | Caliber: sources, as-of, which book, which session, material gaps. | One compact block, one line each — D7. | + +**The top is an answer, not a topic sentence.** It names the stance — +proceed, resize, delay, collect evidence, pick this candidate, no trade, or +"this is what your book now looks like" for a question with no action in it — +and the single reason that decides it. A first sentence that describes what +the answer is about, restates the question, or narrates what was computed has +spent the only position the reader is guaranteed to read. + +**The middle is gated by increment.** Every block must add a NEW +decision-relevant fact or judgment. The test is not "is this true" and not "is +this owed" — under the whitelist era everything printed was both. The test is: +**delete this block; does the decision change?** If it does not, the block is +not shortened, it is deleted. Blocks that survive the gate have no cap, and +this is deliberately not a length rule: an answer that genuinely needs six +increments gets six. + +**Four named bans.** Each was observed live in the #827 or #830 runs, each +survived every rule then on the books, and each is a way for a block to add +volume without adding an increment. The slug after each name is what the +exemplar corpus references it by (§3.5): + +- **A manufactured hypothetical scenario nobody asked for** + (`manufactured_scenario`). Inventing a comparison, a simulation, or a + what-if to carry facts that had nowhere else to go. The owner's verdict on + the observed instance — an all-in-one-name simulation nobody requested — was + that it meant nothing. The user's own question already bounds what is being + decided. +- **A system default restated as insight** (`default_as_insight`). The + engine's own threshold fired, and the answer explains the threshold as though + the reader had learned something about their book. Both arms of the #827 A/B + independently wrote the same sentence about the same default, which is what + obligation discharge looks like from the outside. The user's *own* rule is + the opposite case and is never optional (see §4's D-series and + `rule_effects`). +- **A hedging couplet** (`hedging_couplet`). "This is not a reason to wait, + but…" — a sentence that states a position and withdraws it in the same + breath, leaving the reader with the work of deciding which half was meant. + Say the half you mean. A real uncertainty is a falsifier with a threshold, + not a hedge. +- **The same point in a second form** (`restated_point`). A judgment made in + prose and then again as a bullet, a table row, or a summary line. Restating + is not emphasis; it is the reader paying twice for one increment. This is + D6/D7 seen from the shape side. + +**The bottom floor is a door, not a section.** The route still computes +everything, and the user may ask for any of it at any moment. The answer says +**once** that expansion is available — one short offer, in whatever register +the surface speaks — and does not preview, summarize, or partially deliver the +expansion in advance. Not saying a number this decision does not turn on is +not hiding it; that is the whole distinction #830's deletion rests on. + +**The end block is D7's one floor**, unchanged by this chapter: sources, +as-of, which book, which session, and the material gaps, one line each, +non-narrative, collected once. An answer with nothing on that floor ends at +its judgment (D4). + +### 3.2 Voice + +Write as a senior analyst speaking to their own principal. Direct, concrete, +and already inside the decision: the reader owns the money, has the context, +and asked a real question. Name the call and the number it turns on. Do not +teach the concept behind the number, do not narrate the process that produced +it, do not soften a judgment into a menu, and do not close with a summary the +reader just read. Confidence is expressed as a threshold that would change the +call, never as an adverb and never as a disclaimer. + +### 3.3 Derivation is additive, and empty derivation is the default + +A surface may **add parameters** to this shape: which blocks its middle floor +is allowed to hold, which question set it may ask from, what its end block +must name, what it may not compute at all. A surface may **not** restate, +narrow, re-order, or contradict the shape itself. A surface with no special +need declares nothing and falls through to the default — that is the expected +case, not a gap. + +The one document incarnation is the review card: its keynote plus four blocks +*is* this pyramid rendered as a document, and `output-contract.md` §2 keeps +that structure exactly as it is. Its derivation is recorded there; nothing +about the card changes. + +### 3.4 Registry freeze for shape and length + +**V, D, and C take no new IDs for a shape or length concern.** A style fix has +exactly two lanes now: + +1. **Amend this chapter** — requires an owner ruling, and lands in §8's log. +2. **Add an exemplar or a counter-exemplar** — the day-to-day lane, no ruling + required, and the one that carries most fixes. + +A new V/D/C ID for "answers are too long", "lead with X", or "stop repeating +Y" is refused: five independent phrasings of the answer-first principle is +what this chapter exists to end, and a sixth with an ID on it is still a sixth. +The registries keep their existing IDs and their existing subjects — a +disclosure's *materiality* (D), a citation's *provenance* (C), a lead's +*failure class* (V) — and the historical definitions stay readable where they +are. Where a local answer-shape phrasing was superseded by this chapter, the +surface document says so rather than deleting its own history. + +### 3.5 Exemplars are the spec + +The binding statement of this shape is not the prose above; it is the exemplar +set in `tests/agent/expression-witnesses.json`. Three to five canonical +exemplars per conversational surface, each declaring the one-sentence answer it +leads with and the increment each of its blocks adds, plus counter-exemplars +for the named bans. Every issuer in them is fictional (Widgetron WDGT, +Gridcore GRDC, Fabrion FABR, ACME) and nothing is derived from a user record. + +`tests/agent/check_expression.py` derives E-7 and E-8 from that corpus (§7). +What they decide is exactly two things — that the declared answer really leads, +and that every declared block really adds a distinct increment — and the +corpus records, as an asserted fact, that two of the four bans are +**invisible** to it. A manufactured scenario declares no increment and a +restated point declares one twice, so E-8 reaches both. A system default +explained as insight and a hedging couplet are well-formed blocks carrying +true content; deciding that one is a sermon and the other withdraws its own +position requires reading what the answer means, which is the boundary +`docs/development-guide.md` already draws between a code check and a judge. +This chapter states that boundary rather than implying a gate it does not have. + +## 4. Disclosure relevance and placement (D1–D7) A disclosure is a sentence about the *limits* of what was just said: the book it was measured on, the session it was priced at, the part of the denominator @@ -180,7 +346,7 @@ Four consequences the surfaces below inherit: everything, and the user can ask for any of it at any time. Not saying a number this decision does not turn on is not hiding it. -## 4. Citation and provenance (C1–C4) +## 5. Citation and provenance (C1–C4) | ID | Rule | Verification class | Named oracle | |---|---|---|---| @@ -235,36 +401,44 @@ Already enforced on the question surface and on `--agent-case` claims every surface, and `check_expression.py` E-5 is its check on the conversational ones. -## 5. Surface map +## 6. Surface map -Each surface document keeps its own layout and references this file. None of -them may restate, narrow, or contradict V/D/C. +Each surface document keeps its own layout, derives its answer shape from §3, +and references this file. None of them may restate, narrow, or contradict the +pyramid or V/D/C. "Derivation" is what that surface **adds** to §3; an empty +derivation is valid and is the default. -| Surface | Layout authority | What it keeps | -|---|---|---| -| Review card | [output-contract.md](output-contract.md) | Keynote + four blocks, module prerequisites, which block the footnote ends. | -| `consider` | `references/trade-consequence.md` | Salience order, answer slots, what the payload means. | -| Freeform answers | `references/freeform-answers.md` | Text-first defaults, proportionate production, reusable engine-backed views. | -| No recorded book | `references/decision-framing.md` | Claim boundaries, question heuristics, the strategy-class map, the invitation set. | -| Weekly market read | `references/weekly-market-read.md` | What the prototype reads, and what it may not invoke. | +| Surface | Layout authority | Derivation it adds to §3 | Everything else it keeps | +|---|---|---|---| +| Review card | [output-contract.md](output-contract.md) | The **document incarnation**: keynote + four fixed blocks, in that order, on every committed card. | Module prerequisites, which block the footnote ends. | +| `consider` | `references/trade-consequence.md` | Lead-selection salience order; the answer slots the middle floor may hold; `rule_effects` is never traded away. | What the payload means, the obligation floor. | +| Freeform answers | `references/freeform-answers.md` | None on shape — text-first is a latency default, not a shape. | Proportionate production, reusable engine-backed views. | +| No recorded book | `references/decision-framing.md` | The top sentence is a research-backed baseline when no book exists; the strategy-class map is a middle-floor block set. | Claim boundaries, question heuristics, the invitation set. | +| Weekly market read | `references/weekly-market-read.md` | Its one optional question comes after the complete brief, never before it. | What the prototype reads, and what it may not invoke. | -## 6. Enforcement +## 7. Enforcement | Check | What it observes — and what it does not | Where | |---|---|---| | `check_card.py` S-3 | Review-card layout only: no consecutive caveat paragraphs, none before Block 1, none inside Block 1. It does not govern conversational placement. | `tests/agent/check_card.py` | | `check_expression.py` E-5 | C4 only: no engine payload token reaches a conversational answer. E-1–E-4 were retired by #825 because formatting is not evidence of relevance or clarity. | `tests/agent/check_expression.py` | | `check_expression.py` E-6 | D7's bottom floor only: no machine anchor is rendered. Matched by shape (a long hex run), not by vocabulary, so it survives a prefix rename without importing the engine. It says nothing about which floor an owed fact landed on. | `tests/agent/check_expression.py` | +| `check_expression.py` E-7 | §3's top floor, on an exemplar: the scene's declared one-sentence answer must actually appear in the answer's opening block. It decides *placement of a declared core*, never whether that core is the right call. | `tests/agent/check_expression.py` | +| `check_expression.py` E-8 | §3's increment gate, on an exemplar: every declared block's text occurs in the answer in declared order, and every block declares a distinct, non-empty increment. It decides *distinctness*, never whether an increment was worth having. | `tests/agent/check_expression.py` | | `ux_receipt` delivery evidence | The same rule where the frozen value and the presented text are both in hand: a declared `machine_state` value appearing in the answer refuses the evidence rather than counting it. Silent on a challenge block that predates the key. | `skills/fomo-kernel/tools/ux_receipt.py` | | `check_voice.py` | V1–V9 witness classification. | `tests/agent/check_voice.py` | | `answer_provenance` | C1/C2/C4 on a structured `--agent-case`, and the coverage a case may not leave uncited — including the extent of an illegible book, not only that it is one. | `skills/fomo-kernel/engine/answer_provenance.py` | -| `test_expression_contract.py` | That both registries are complete, every surface routes here, D1–D6 honestly declare instruction-only verification, the C4 blacklist remains schema-derived, and the `consider` obligation floor stays smaller than the whole computed inventory. | `tests/test_expression_contract.py` | +| `test_expression_contract.py` | That both registries are complete, every surface routes here **and declares its §3 derivation**, that no surface still carries a local answer-first phrasing, D1–D6 honestly declare instruction-only verification, the C4 blacklist remains schema-derived, and the `consider` obligation floor stays smaller than the whole computed inventory. | `tests/test_expression_contract.py` | **None of these runs against a live answer.** `check_expression.py` proves only -the exact C4 and D7-floor properties it can decide. D1–D6, and D7's other three -floors, are evaluated by reading the answer in context; pretending a regex -covered them was the constraint failure #825 removed. Nothing sits between the -model and the user. +the exact C4, D7-floor, and exemplar-corpus properties it can decide. D1–D6, +D7's other three floors, and two of §3's four named bans — a system default +restated as insight, and the hedging couplet — are evaluated by reading the +answer in context; pretending a regex covered them was the constraint failure +#825 removed. The corpus asserts that gap rather than hiding it: a +counter-exemplar for either of those two bans must **pass** every mechanical +assertion, so the day someone builds a real oracle for one of them, that +assertion is what tells them the coverage boundary moved. Nothing sits between the model and the user. The delivery half is observed the same way every other instruction-tier rule in this repository is: by owner-live dogfood and by the QA receipt @@ -272,7 +446,7 @@ in this repository is: by owner-live dogfood and by the QA receipt passing loader is implementation evidence and this file says so rather than letting a green suite read as a governed output. -## 7. Ruling log +## 8. Ruling log | Date | Ruling | |---|---| @@ -281,3 +455,4 @@ letting a green suite read as a governed output. | 2026-08-14 | Line cap set at five (D5) from a measured four-line worst case, with merging — never dropping — as the remedy, so a cap can never become an argument for omitting an owed fact. | | 2026-08-19 | Issue #825 retires the product-wide block, prefix, and line-cap template plus E-1–E-4. Those checks proved formatting, not whether a limitation mattered. The review card keeps its footnote as local layout; conversational surfaces use relevance-driven placement. | | 2026-08-20 | Owner ruling ([#830](https://github.com/atomchung/fomo-kernel/issues/830)): the product is too verbose, and the fix is deletion rather than a reading budget — anything whose must-have reason cannot be stated is cut. `consider`'s obligation list splits into owed / available / never-rendered, and D7 makes volume *distribution* expression's business, which §2 had disclaimed and nothing else had claimed. The reading-budget rule proposed as V10 is demoted to a backstop and is not adopted here: a length cap is what #827 had just deleted, and re-adding one would have priced the symptom instead of removing the cause. | +| 2026-08-21 | Owner ruling ([#832](https://github.com/atomchung/fomo-kernel/issues/832)): **one communication method.** §3 becomes the mother law — one pyramid, mandatory on every surface — and the five independent local phrasings of answer-first (the card keynote, V1, SKILL.md's reader-order paragraph, `decision-framing.md`'s bounded-value lead, `weekly-market-read.md`'s value-first, plus `trade-consequence.md`'s reader-question-chain section, the sixth the audit found) are replaced by derivation references. Derivation is **additive only** and an empty derivation is the default. The increment gate and four named bans are encoded; V/D/C take no new IDs for a shape or length concern, so a future style fix is either an owner amendment to §3 or an exemplar — never a sixth phrasing with an ID on it. The card's structure is unchanged: it is the document incarnation of the same pyramid. Deliberately **not** adopted, again: any character-count cap (#543's ceiling, deleted by #827, stays deleted) — the shape is positive, and length is its consequence rather than its rule. Integrity gates (engine-owned numbers, provenance, canonical writes, execution truth, privacy) are untouched, and the #829 unselected-write finding stays open and out of scope. | diff --git a/docs/maintainer-guide.md b/docs/maintainer-guide.md index 3f2bdd2..79ae843 100644 --- a/docs/maintainer-guide.md +++ b/docs/maintainer-guide.md @@ -140,14 +140,19 @@ a later row explicitly supersedes an earlier answer-shape rule, the later row is current authority and the earlier row remains historical context only. In particular, #827 supersedes #543's effort/closed-chart ceiling and #674's fixed refusal shape while preserving their latency, numeric-grounding, and state -integrity findings, and #830 supersedes the whitelist-era answer obligations +integrity findings, #830 supersedes the whitelist-era answer obligations (#479 Wave B cut 2, #823, #825) while preserving every integrity gate they -built. +built, and #832 supersedes the **answer-shape phrasing** in all five surfaces +that carried one (#823's per-surface layout classification, the card keynote, +V1, SKILL.md's reader-order paragraph, `decision-framing.md`'s bounded-value +lead, `weekly-market-read.md`'s value-first, plus `trade-consequence.md`'s +reader-question-chain) while leaving those rows' every other ruling intact. | Fact | Surfaces that must stay synchronized | |---|---| | Output structure & language | `docs/output-contract.md` (single authority on section order) and `docs/output-language.md` (locale contract) ↔ `card_renderer.py` ↔ `references/card-policy.md` / `card-spec.md` (subordinated: wording and in-block ranking only) | | How the product speaks, on every surface (#823, #825) | `docs/expression-contract.md` is the single authority for expression — voice routes to `docs/output-voice.md`, disclosure relevance is D1–D6, and citation/provenance is C1–C4. Its surfaces are `docs/output-contract.md` §4 ↔ `references/trade-consequence.md` ↔ `references/freeform-answers.md` ↔ `references/decision-framing.md` ↔ `references/weekly-market-read.md`. The card may keep its footnote as local layout, but conversational surfaces have no mandatory tail block, prefix, or line cap. `tests/agent/check_expression.py` now enforces only E-5/C4 (no engine-token leaks); #825 retired E-1–E-4 because block position, marker syntax, line count, and literal deduplication did not prove relevance or clarity. `tests/test_expression_contract.py` keeps registry/routing checks and pins that D1–D6 honestly declare instruction-only verification. | +| One communication method: the answer pyramid (#832, 2026-08-21) | `docs/expression-contract.md` §3 is the **mother law** of answer shape and the only statement of it. Its readers, each of which now carries a *derivation* rather than a phrasing: `docs/output-contract.md` §2 (the card is the document incarnation — keynote = top floor, three middle blocks = the middle, the Block-1 footnote = the end block; the card's own structure is unchanged) ↔ `docs/output-voice.md` (V1 keeps its ID as the **failure class** and stops being a second statement of the rule) ↔ `skills/fomo-kernel/SKILL.md` "Shape of the answer" (the always-on projection, explicitly labelled as one) ↔ `references/trade-consequence.md` (adds exactly two parameters: lead selection and answer slots) ↔ `references/decision-framing.md` (adds two: the baseline is the top sentence with no book, the strategy-class map is a middle-floor block set) ↔ `references/weekly-market-read.md` (adds one: its optional question comes after the complete brief) ↔ `references/freeform-answers.md` (**adds nothing — an empty derivation is valid and is the default**) ↔ `tests/agent/expression-witnesses.json` ↔ `tests/agent/check_expression.py` (E-7/E-8) ↔ `tests/test_expression_contract.py` ↔ `tests/test_research_priors.py` (which pinned the no-book phrasing by literal and now pins the block order and the derivation instead). Four rules. **Derivation is additive-only**: a surface may add which blocks its middle may hold, which questions it may ask, what its end block must name, or what it may not compute — it may not restate, narrow, re-order, or contradict the shape. **The registries are frozen for shape and length**: V, D, and C take **no new ID** for "answers are too long", "lead with X", or "stop repeating Y", because five independently-worded statements of answer-first is the disease and a sixth with an ID on it is still a sixth; V10 stays unallocated (proposed in #830, demoted there, refused here). **A style fix has two lanes and only two**: amend §3 (owner ruling required, logged in §8) or add an exemplar/counter-exemplar to the witness corpus (day-to-day, no ruling). **The exemplars are the spec, and the oracle is honest about its half**: E-7 decides that a scene's declared one-sentence answer really leads, E-8 that every declared block really adds a distinct increment; the manufactured-scenario and hedging-couplet bans have **no** mechanical oracle, and the corpus asserts that by requiring their counter-exemplars to *pass* every assertion — the day one of them can be caught, that assertion is what says the boundary moved. Why the row exists at all: #830's post-merge rerun (PR #831) fixed composition and moved first-answer length by less than 5%, which is the evidence that deleting obligations without a positive shape norm only vacates space for discretionary elaboration. **Not adopted, again**: any character-count cap — #543's ceiling stayed deleted, and length is the shape's consequence, never its rule. Integrity gates (engine-owned numbers, provenance, canonical writes, execution truth, privacy) and #829 are untouched and out of scope. | | Runtime behavior | engine ↔ `SKILL.md` and routed flows/references ↔ `docs/eval-design.md` ↔ `evals/EVALS.md` | | Demo card values | English README ↔ English demo HTML/image; Traditional Chinese README ↔ Traditional Chinese demo HTML/image. Values must match; only wording differs. | | GTM documentation | `README.md` is the English default; `README.zh-TW.md` is the complete Traditional Chinese counterpart. Keep language links and substantive product claims synchronized. | diff --git a/docs/output-contract.md b/docs/output-contract.md index b003466..7447353 100644 --- a/docs/output-contract.md +++ b/docs/output-contract.md @@ -30,6 +30,17 @@ ## 2. Canonical structure: keynote + four blocks +**This structure is the document incarnation of the answer pyramid** +([expression-contract.md](expression-contract.md) §3, owner ruling +2026-08-21 — #832). The keynote is the pyramid's top floor rendered as a +document, the three middle blocks are its increment-gated middle, and Block 1's +footnote is its end block. Nothing about the card changes: the derivation is +recorded so the card stops being read as an independent statement of +answer-first, which is what five surfaces each saying it their own way had +already cost. The keynote's own rules below — one sentence, the period's most +important judgment, the review window on its own line — are this surface's +**added parameters**, not a second answer-shape rule. + Every committed review card renders, in this order: | # | Block | Content | Demo-card anchor | diff --git a/docs/output-voice.md b/docs/output-voice.md index 33ed88e..7604f35 100644 --- a/docs/output-voice.md +++ b/docs/output-voice.md @@ -17,8 +17,18 @@ Voice is one of three expression registries. [expression-contract.md](expression-contract.md) routes to this file for V1–V9 and owns the other two: **where a disclosure goes, and on which floor** (D1–D7) and **how a claim states where it came from** (C1–C4). A rule about placement or citation belongs -there, not here; a rule about what an answer leads with, refuses to -manufacture, or stops at belongs here. +there, not here; a rule about what an answer refuses to manufacture or stops at +belongs here. + +**The shape of an answer is not this file's** (#832, 2026-08-21). The +[expression contract](expression-contract.md)'s §3 mother chapter is the one +statement of it — one-sentence answer on top, an increment-gated middle, the +rest of the inventory behind a single offer, one caliber block at the end — and +V1 derives from it rather than stating it a second time. **This registry takes +no new ID for a shape or a length concern**: a style fix is either an owner +amendment to §3 or a new exemplar in +`tests/agent/expression-witnesses.json`. V10, proposed as a reading budget in +#830 and demoted there, stays unallocated for the same reason. Phase 1 integrates and proves this authority only on `consider` and no-book decision framing. That limited proof does not exempt other surfaces; it avoids @@ -44,9 +54,16 @@ deterministic product truth. ## Rules -- **V1 — decision value before boundary.** Lead with the supported decision - tension or completion, not an error, process description, disclaimer, or - generic limitation. +- **V1 — decision value before boundary.** What an answer leads with is the + pyramid's top floor and + [expression-contract.md](expression-contract.md) §3 owns it. V1 keeps its ID + as the **failure class**: an answer that opens on an error, a process + description, a disclaimer, or a generic limitation instead of the supported + decision is classified V1 by the witness oracle and by every cross-host run + recorded under that ID. Its historical definition — "lead with the supported + decision tension or completion" — is superseded by §3 as a *statement of the + rule*, and preserved here because the fixtures and rulings that cite V1 are + about this failure. - **V2 — nearest useful completion.** If a requested action or calculation is unavailable or out of scope, complete the closest allowed reasoning task rather than stopping at the boundary. diff --git a/skills/fomo-kernel/SKILL.md b/skills/fomo-kernel/SKILL.md index 92c4738..758d104 100644 --- a/skills/fomo-kernel/SKILL.md +++ b/skills/fomo-kernel/SKILL.md @@ -9,7 +9,7 @@ Use relevant evidence and the recorded book when portfolio consequences matter; ## Answer a live decision -Use `consider` when the user supplies a trade premise and asks what it does to a recorded book. It is the deterministic portfolio-consequence path, not a prerequisite for company research, candidate discovery, or a non-portfolio recommendation. +Use `consider` when the user supplies a trade premise and asks what it does to a recorded book. It is the deterministic portfolio-consequence path, never a prerequisite for research, discovery, or a non-portfolio recommendation. ```bash cd skills/fomo-kernel @@ -20,7 +20,7 @@ A premise needs a `ticker`, a `side`, and one of `qty` or `notional`. Everything Pass `--language` as the tag the user is writing in; an unsupported tag falls back to `en`. Keep conversing in their language and never hand-translate engine copy. -First run only: `pip install -r requirements.txt`, then `python3 engine/review.py doctor`. The engine fail-soft degrades without its optional dependencies — silently dropping current prices, P&L, alpha/beta, and market context — so verify once rather than discovering it inside an answer. +First run only: `pip install -r requirements.txt`, then `python3 engine/review.py doctor`. The engine fail-soft degrades without its optional dependencies — silently dropping current prices and market context — so verify once rather than mid-answer. ## The response is the contract @@ -36,15 +36,15 @@ Read portfolio consequence from that payload; never recompute or fill its gaps. ## Research only what could change the recommendation -The engine computes portfolio consequence; it is not a company-research service. Do not fetch a standing market packet on every call. Look up current price, recent movement, valuation, an event, or operating evidence only when that fact is material to the user's question or could change the recommendation. A found event never becomes the user's motive until they confirm it is. `references/market-lookup.md` owns the bounded lookup and provenance contract. +The engine computes portfolio consequence; it is not a company-research service. Look up current price, recent movement, valuation, an event, or operating evidence only when that fact is material to the user's question or could change the recommendation. A found event never becomes the user's motive until they confirm it is. `references/market-lookup.md` owns the bounded lookup and provenance contract. ## Shape of the answer -Answer in the reader's own order — what they asked, the answer, why, what would overturn it, what to do — not the payload's field order. Say what this decision needs, not what exists. Open on the stance and the reason that decides it: proceed, resize, delay, collect evidence, choose one candidate, or no trade. Carry the one or two numbers that would flip it, and every `rule_effects` entry, which is never optional. Give a directional call a falsifier — that is the counter-case, and it needs no section. Ask only decision-changing questions, then stop. +One shape, every answer (`../../docs/expression-contract.md` §3 owns it; this is its projection, not a second wording). **A fact lives on exactly one floor, and twice is a bug.** *Top:* one sentence — the stance and the reason that decides it (proceed, resize, delay, collect evidence, choose one candidate, no trade). *Middle:* only blocks that add a new decision-relevant fact or judgment — delete one; if the decision does not change, delete it. There live the numbers that would flip the call, every `rule_effects` entry (never optional), a truth-critical denominator, unit, or pricing set beside its number, and a falsifier on any directional call — the counter-case needs no section. *Bottom:* the rest of the inventory stays in the data layer; say once you can expand it. *End:* one compact block for other material limitations; machine anchors and engine narration nowhere. -**A fact lives on exactly one floor; twice is a bug** — deciding facts in the body, a truth-critical denominator, unit, or pricing set beside its number, other material limitations in one compact end block, machine anchors and engine narration nowhere. `references/trade-consequence.md` holds the rest. +Never manufacture a scenario nobody asked for, restate a system default as insight, hedge in couplets, or make one point twice. Ask only decision-changing questions, then stop. `references/trade-consequence.md` holds the rest. -Label your thesis, valuation, timing, forecast, recommendation, ranking, or selection as judgment, separate from engine facts. Give a target or forecast's material assumptions and uncertainty; never disguise it as fact or certainty. Never claim what the user did or will do. +Label judgment — thesis, valuation, timing, forecast, recommendation, ranking, selection — separate from engine facts. Give a target or forecast's material assumptions and uncertainty; never disguise it as fact or certainty. Never claim what the user did or will do. **Candidate discovery and comparison.** For an explicit search, report universe, filters, as-of point, material exclusions, and coverage limits; never imply @@ -59,7 +59,7 @@ evaluation row. - **Unpriced instruments.** The payload says how to return them. Read closes from the publisher's page, transcribe the `references/price-feed.md` envelope, and rerun with `--prices `. If none are published, `--prices-unavailable ''` refuses only the current-value portfolio consequence; still give supported non-portfolio judgment. Never invent, interpolate, or recall a price; missing is not delisted or zero. - **No recorded book.** `consider` fails closed for book-derived claims. Continue with supported research and judgment, and frame the decision under `references/decision-framing.md`; do not manufacture portfolio precision or persist the conversation. -A refusal does not end the turn. You still owe the judgment that holds without the numbers the engine would not compute — say plainly what could not be checked and name what would unblock it, because the user's next move is to close that gap. Never present a degraded number as if it were the real one: a forward-looking decision is refused rather than answered on cost weights precisely because cost weights can invert which position is the largest. +A refusal does not end the turn. You still owe the judgment that holds without the numbers the engine would not compute — say plainly what could not be checked and name what would unblock it. Never present a degraded number as if it were the real one: a forward-looking decision is refused rather than answered on cost weights, which can invert which position is the largest. ## After the answer @@ -73,7 +73,7 @@ python3 engine/review.py consider --resolve --decision acted|dec ## Other jobs -Reach for these when the user asks for them. None of them routes an ordinary decision. +Reach for these when the user asks. None routes an ordinary decision. | The user wants | Do this | |---|---| @@ -84,4 +84,4 @@ Reach for these when the user asks for them. None of them routes an ordinary dec | To continue after an interruption | `python3 engine/review.py resume` — never refetch prices mid-session | | A failed projection repaired | `python3 engine/review.py repair-projections` | -A simple ad hoc question defaults to a fast, direct text answer. Use relevant research, multiple tools, or a visual when the user asks or when it materially improves the decision; keep the work proportionate and report material coverage limits. +A simple ad hoc question defaults to a fast, direct text answer; scale research, tools, and visuals to decision value and report material coverage limits (`references/freeform-answers.md`). diff --git a/skills/fomo-kernel/engine/evaluation_challenge.py b/skills/fomo-kernel/engine/evaluation_challenge.py index d17aff7..c1c076c 100644 --- a/skills/fomo-kernel/engine/evaluation_challenge.py +++ b/skills/fomo-kernel/engine/evaluation_challenge.py @@ -186,8 +186,8 @@ # # `concentration` and `cash` left this tuple in #830 and now live in # `MAY_STATE_TOPICS`. The order that remains is a dependency order, never a -# reading order: `references/trade-consequence.md` states the reader's own -# question chain the answer is arranged by, and the deciding fact opens it. +# reading order: `docs/expression-contract.md` §3 owns the shape the answer is +# arranged in, and the deciding fact opens it. TOPICS = ("basis", "price_basis", "position", "rule_collision", "disclosure", "excluded_holding", "out_of_scope") diff --git a/skills/fomo-kernel/references/decision-framing.md b/skills/fomo-kernel/references/decision-framing.md index 55c76a9..5ccea92 100644 --- a/skills/fomo-kernel/references/decision-framing.md +++ b/skills/fomo-kernel/references/decision-framing.md @@ -24,10 +24,19 @@ when it could change the recommendation or unlock the portfolio claim. ## Voice and expression authority Apply the global [expression contract](../../../docs/expression-contract.md): -voice through the [output-voice contract](../../../docs/output-voice.md) -(V1–V9), disclosure relevance and placement through D1–D7, provenance labelling through -C1–C4. They own universal output semantics; this reference owns the no-book -facts, questions, and route order below. +the answer's shape through its §3 mother chapter, voice through the +[output-voice contract](../../../docs/output-voice.md) (V1–V9), disclosure +relevance and placement through D1–D7, provenance labelling through C1–C4. +They own universal output semantics; this reference owns the no-book facts, +questions, and route order below. + +**This route's derivation from §3, and nothing more** (#832): with no book, +the pyramid's top sentence is a *research-backed baseline* rather than a +computed consequence, and the strategy-class map below is a middle-floor block +set. Both are stated once, in "Research-aware strategy framing". Everything +else about the shape — that the top is one sentence, that every block must add +a new decision-relevant fact, that the rest of the inventory waits behind one +offer — is §3's, and this file no longer says it a second time. ## What the answer is @@ -45,11 +54,12 @@ A useful framing may carry: When a user asks for a strategy before they have a book, do not make them invent an exit philosophy before supplying the bounded value available now. -For a simple strategy question, lead with the bounded value already supported: +This route's block order — the parameter it adds to the pyramid, whose top +floor is already the answer: ```text -research-backed baseline -→ applicable strategy-class map +research-backed baseline (the top sentence, when no book exists) +→ applicable strategy-class map (middle floor) → any question whose answer could change the recommendation ``` diff --git a/skills/fomo-kernel/references/freeform-answers.md b/skills/fomo-kernel/references/freeform-answers.md index 0b796ea..9c5c124 100644 --- a/skills/fomo-kernel/references/freeform-answers.md +++ b/skills/fomo-kernel/references/freeform-answers.md @@ -163,6 +163,14 @@ every limitation. Keep truth-critical denominator, unit, or pricing-set qualifiers inline. Include other limitations only when they materially qualify the answer; no marker or numeric line cap is required. +**This route's derivation from §3 is empty, and that is the expected case** +(#832). A freeform answer takes the pyramid exactly as it stands — one sentence +on top, an increment-gated middle, the rest of the inventory behind one offer, +one caliber block at the end — and adds no parameter of its own. Rule 1's +text-first default is a **latency** preference about how much work to do before +answering; it is not a shape, it never was, and reading it as one is how a +surface acquires a second answer-shape rule. + Since #830, *where* they go is a rule rather than a free choice. A fact lives on exactly one floor (D7): the facts that decide the call open the body, a truth-critical qualifier stays beside its number, and everything else material diff --git a/skills/fomo-kernel/references/trade-consequence.md b/skills/fomo-kernel/references/trade-consequence.md index 88ac7da..3180a18 100644 --- a/skills/fomo-kernel/references/trade-consequence.md +++ b/skills/fomo-kernel/references/trade-consequence.md @@ -288,7 +288,7 @@ bullet each. `must_state`'s order is the order the facts depend on each other the price session next because those numbers are measured at one, the disclosures last because they qualify what precedes them). It is a dependency order and never a reading order: [the answer's own order](#answer-shape) is the -reader's question chain, and the fact that decides the call opens it. +pyramid's, and the fact that decides the call opens it. ### The three lists, and why each fact is on the one it is on (#830) @@ -355,35 +355,40 @@ Under maintainer QA, delivery of these obligations is proven rather than assumed ## Route-specific synthesis Apply the global [expression contract](../../../docs/expression-contract.md): -voice through the [output-voice contract](../../../docs/output-voice.md) -(V1–V9), disclosure relevance and placement through D1–D7, provenance labelling through -C1–C4. Those own how this answer speaks; this section owns only the `consider` -route's salience facts and answer slots. +the answer's shape through its §3 mother chapter, voice through the +[output-voice contract](../../../docs/output-voice.md) (V1–V9), disclosure +relevance and placement through D1–D7, provenance labelling through C1–C4. +Those own how this answer speaks; this section owns only the `consider` route's +salience facts and answer slots. The challenge block is the factual floor. The shape below selects the decision-relevant facts without turning available ones into standing copy. -### The reader's question chain +### Derivation from the pyramid -Arrange the answer in the order the reader would ask it, never in the payload's -field order: +The answer's shape is [expression contract §3](../../../docs/expression-contract.md)'s +— one sentence on top, an increment-gated middle, the rest of the inventory +behind a single offer, one caliber block at the end — and this file states it +nowhere else (#832; before that, the reader's-question-chain paragraph here was +one of six independent phrasings of the same idea). -**what you asked → the answer → why → what would overturn it → what you'd do.** - -That is one storyline, which is V4's one lead tension applied to the whole -answer rather than only to its opening. Payload order is a dependency order -computed for a machine; reading it aloud is how an answer comes to open on the -book's provenance and reach the recommendation in its last paragraph. +Two parameters this route adds, and only these two: **which fact wins the top +sentence** (lead selection, below) and **which blocks the middle floor may +hold** (answer slots, below). Payload order is a dependency order computed for +a machine; reading it aloud is how an answer comes to open on the book's +provenance and reach the recommendation in its last paragraph, which is what +the pyramid exists to prevent. ### Lead selection Unless a truth-critical disclosure changes how an earlier item can be understood, salience runs: -1. **A user-authored rule collision** — `rule_effect` of `new_breach` or `worsened_existing_breach`. The user wrote that line themselves; this trade crossing it or digging further into it outranks everything else. -2. **The largest non-obvious portfolio consequence** — weight, concentration or driver overlap, cash. *Non-obvious* is load-bearing: the user already knows they hold the position and that the price fell. What they cannot see from where they sit is what the trade does to the whole book's shape. -3. **The decision-context read** — whether `why_now` looks like a real evidence delta or a price move wearing one, labelled as your judgment. [market-lookup.md](market-lookup.md) governs verifying it. -4. **Routine basis and unchecked boundaries** — include them when they materially qualify or could reverse the recommendation. Do not append a standard tail merely because a field exists. +1. **A funding shortfall** — a negative post-trade cash balance (#778). It outranks everything else when it occurs, because it is not a portfolio consequence at all: it says this trade cannot be done out of the recorded book, so the user is either funding it from somewhere the engine cannot see or selling something to do it. Lead with **that decision**, not with the balance and its weight. Those two numbers are owed and they are support, not the answer — an answer that recites them and then explains what the engine can and cannot see has led with the boundary, which is the exact defect #778 recorded. The boundary is still stated, once, beside the claim it qualifies (D2/D6): the engine sees the recorded book and no other account. Never assume the user has other cash, and never assume they do not. +2. **A user-authored rule collision** — `rule_effect` of `new_breach` or `worsened_existing_breach`. The user wrote that line themselves; this trade crossing it or digging further into it outranks everything else below. +3. **The largest non-obvious portfolio consequence** — weight, concentration or driver overlap, cash. *Non-obvious* is load-bearing: the user already knows they hold the position and that the price fell. What they cannot see from where they sit is what the trade does to the whole book's shape. +4. **The decision-context read** — whether `why_now` looks like a real evidence delta or a price move wearing one, labelled as your judgment. [market-lookup.md](market-lookup.md) governs verifying it. +5. **Routine basis and unchecked boundaries** — include them when they materially qualify or could reverse the recommendation. Do not append a standard tail merely because a field exists. Special cases: `improved_but_still_over` and `resolved_existing_breach` are improvements to an already-broken line, never framed as a new breach — an improvement that leaves the line crossed still leads with both truths, and one that clears it is worth saying out loud rather than passing in silence. A `partial_book` or missing-FX denominator qualifies every affected percentage in the same sentence — it is the textbook truth-critical qualifier (expression contract D2), because the number means something different without it. A stale or cost-basis book leads only when it makes the apparent consequence unreliable enough to change the decision; otherwise include it only when it materially qualifies the recommendation. With no collision, lead with the largest changed consequence; with no material change, say that the supported dimensions show little change and name only what stays materially unchecked — never convert "not measured" into "no risk". diff --git a/skills/fomo-kernel/references/weekly-market-read.md b/skills/fomo-kernel/references/weekly-market-read.md index 896b525..4114cb0 100644 --- a/skills/fomo-kernel/references/weekly-market-read.md +++ b/skills/fomo-kernel/references/weekly-market-read.md @@ -26,15 +26,20 @@ must follow `market-lookup.md`: one triggered packet maximum, source/as-of on every public fact, and never infer the user's motive. How that brief is said is not this file's to decide: apply the global -[expression contract](../../../docs/expression-contract.md) — voice V1–V9, -disclosure relevance and placement D1–D7, provenance labelling C1–C4. Source and as-of on a -public fact are C2; the labelled judgment risk is C1. This file owns only what -the read may compute and what it must refuse. +[expression contract](../../../docs/expression-contract.md) — the answer's +shape through its §3 mother chapter, voice V1–V9, disclosure relevance and +placement D1–D7, provenance labelling C1–C4. Source and as-of on a public fact +are C2; the labelled judgment risk is C1. This file owns only what the read may +compute and what it must refuse. -The host shows value first. The first response has `optional_question.selected` -as `null`, then may ask its one optional question after the complete brief. -When the user skips, stop: the shown brief is already complete. Only on an -answer, rerun the same read-only command with its offered `--focus` value; -that second response has the selected value and a visibly different -current-session watch, without another question. Persistence is outside this -prototype. +**This route's derivation from §3, and nothing more** (#832): the one optional +question comes *after* the complete brief, never before it. That is a sequencing +parameter on a read whose whole first response is a brief; the reason a brief +leads at all is §3's top floor, which this file no longer restates. + +The first response has `optional_question.selected` as `null`, then may ask its +one optional question. When the user skips, stop: the shown brief is already +complete. Only on an answer, rerun the same read-only command with its offered +`--focus` value; that second response has the selected value and a visibly +different current-session watch, without another question. Persistence is +outside this prototype. diff --git a/tests/agent/check_expression.py b/tests/agent/check_expression.py index d883453..41fa5a0 100644 --- a/tests/agent/check_expression.py +++ b/tests/agent/check_expression.py @@ -1,33 +1,54 @@ #!/usr/bin/env python3 # -*- coding: utf-8 -*- -"""Check C4 and D7's machine-anchor floor of the conversational expression -contract (offline, deterministic). - -Issue #825 retired E-1 through E-4. They classified block position, a literal -prefix, line count, and exact-string repetition; none could tell whether a -limitation mattered or where it read clearly. E-5 remains because leaking an -engine payload token is an exact, deterministic product defect. - -E-6 (#830) is the second such defect: rendering a machine anchor. D7 says a -fact lives on exactly one floor, and the bottom floor is "payload only, -never rendered" — `basis.state_version`, a content hash whose only reader is -a mechanical QA comparison. E-6 is the half of D7 a regex can honestly -decide; where the disclosure lands when it *is* owed remains a judgment D1 -and D2 govern. +"""Check the deterministic half of the conversational expression contract +(offline, no dependencies). + +Two of these assertions read an answer's text alone; two read an exemplar +scene, because what they decide is a relationship between an answer and the +shape its author declared for it. + +**Text-only.** E-5 fails an answer that says an engine payload token at the +user (C4). E-6 (#830) fails one that renders a machine anchor: D7 says a fact +lives on exactly one floor, and the bottom floor is "payload only, never +rendered" — `basis.state_version`, a content hash whose only reader is a +mechanical QA comparison. Issue #825 retired E-1 through E-4, which classified +block position, a literal prefix, line count, and exact-string repetition; none +of them could tell whether a limitation mattered or where it read clearly. + +**Exemplar-level (#832).** The answer pyramid — `docs/expression-contract.md` +§3 — is stated bindingly by the exemplar corpus rather than by prose, so its +oracle reads the corpus. E-7 fails a scene whose declared one-sentence answer +(`core`) does not appear in its answer's opening block: the top floor is the +one position the reader is guaranteed to read, and an answer that spends it on +anything else has no core, wherever the rest of it went. E-8 fails a scene +whose declared blocks are not all present in the answer in declared order, or +which declares no increment for a block, or which declares the same increment +twice — the increment gate's mechanical half. + +**What E-7 and E-8 do not decide.** Whether the declared core is the *right* +call, and whether a declared increment was *worth* having, are read, not +matched. Two of §3's four named bans have no assertion at all — a system +default restated as insight, and a hedging couplet — and the corpus asserts +that gap instead of hiding it: a `counter` scene names the ban it violates and +must **pass** every assertion here, so the day an honest oracle for one of them +exists, that scene is what says the coverage boundary moved. A declared +increment is the author's claim about their own scene, exactly as `fails` is; +what this file verifies is that the declaration is faithful to the text (the +block really is there, at that point, exactly once) and internally coherent. The token blacklist for E-5 is READ FROM THE SCHEMAS, never transcribed here. `skills/fomo-kernel/schemas/*.schema.json` already enumerate the engine's own vocabulary, so a new disclosure key or rule effect is covered the day it is added rather than the day someone remembers this file. Hand-mirroring it is what docs/maintainer-guide.md forbids and what borrowed enumerations go stale -from. +from. The ban registry is read the same way, from the fixture itself. Run: - python3 tests/agent/check_expression.py - python3 tests/agent/check_expression.py # check the witness fixture + python3 tests/agent/check_expression.py # E-5 and E-6 only + python3 tests/agent/check_expression.py # the whole corpus Import: from check_expression import check_expression, check_fixture - findings = check_expression(answer_text) # list[Finding] + findings = check_expression(answer_text) # list[Finding], text-only """ from __future__ import annotations @@ -65,7 +86,20 @@ _MACHINE_ANCHOR = re.compile(r"(? Finding: def check_expression(text: str, tokens=None) -> list: - """The deterministic expression assertions against one answer.""" + """The text-only expression assertions against one answer.""" tokens = internal_tokens() if tokens is None else tokens return [_e5_no_internal_tokens(text, tokens), _e6_no_machine_anchor(text)] +# ──────────────────── the pyramid's two decidable halves ──────────────────── + +def _e7_the_core_leads(scene) -> Finding: + """§3's top floor. The scene declares the one sentence that answers the + question; this checks the answer actually opens on it. An opening block is + everything before the first blank line, because that is what a reader gets + without scrolling on every surface this product speaks on.""" + label = "§3 top floor: the declared one-sentence answer leads" + answer = scene.get("answer") or "" + core = scene.get("core") + if not isinstance(core, str) or not core.strip(): + return Finding("E-7", False, label, "the scene declares no core") + if core not in answer: + return Finding("E-7", False, label, "the declared core is nowhere in the answer") + if core not in answer.split("\n\n", 1)[0]: + return Finding("E-7", False, label, + "the declared core is in the answer but not in its opening block") + return Finding("E-7", True, label) + + +def _e8_every_block_adds_something_new(scene) -> Finding: + """§3's increment gate, on the half a match can decide: the declared + decomposition is faithful (each block's text is in the answer, at or after + the previous one) and every block names a distinct increment. A block with + no increment is a manufactured carrier; two blocks with one increment are + the same point in a second form.""" + label = "§3 increment gate: each block is in place and adds a distinct increment" + answer = scene.get("answer") or "" + blocks = scene.get("blocks") + if not isinstance(blocks, list) or not blocks: + return Finding("E-8", False, label, "the scene declares no blocks") + cursor, seen, problems = 0, {}, [] + for index, block in enumerate(blocks): + if not isinstance(block, dict): + problems.append(f"block {index} is not an object") + continue + text = block.get("text") + if not isinstance(text, str) or not text.strip(): + problems.append(f"block {index} declares no text") + continue + position = answer.find(text, cursor) + if position < 0: + problems.append( + f"block {index} is not in the answer at or after the block before it") + continue + cursor = position + len(text) + adds = block.get("adds") + if not isinstance(adds, str) or not adds.strip(): + problems.append(f"block {index} declares no increment") + continue + key = " ".join(adds.split()).casefold() + if key in seen: + problems.append(f"block {index} repeats the increment of block {seen[key]}") + else: + seen[key] = index + return Finding("E-8", not problems, label, "; ".join(problems[:3])) + + +def check_exemplar(scene) -> list: + """The assertions that need the scene's declared shape, not only its text.""" + return [_e7_the_core_leads(scene), _e8_every_block_adds_something_new(scene)] + + # ─────────────────────────── witness fixture ─────────────────────────── def check_fixture(path: pathlib.Path = DEFAULT_FIXTURE) -> list: - """Classify the synthetic witnesses, the way `check_voice.py` does for - V1–V9: a positive scene must produce no finding, and a negative scene - must fail exactly the assertion it declares — no more, so a witness - cannot pass by being broken in several ways at once.""" + """Classify the exemplar corpus. + + A **positive** exemplar must produce no finding. A **negative** one must + fail exactly the assertion it declares — no more, so a witness cannot pass + by being broken in several ways at once. A **counter** exemplar names a ban + nothing here can decide and must pass every assertion; that is the + coverage boundary stated as a test rather than as a promise. + + Three corpus-level properties beyond the per-scene verdicts: every + conversational surface carries §3.5's three to five positive exemplars, + every declared ban is demonstrated by at least one scene, and every + assertion has a negative witness.""" try: fixture = json.loads(path.read_text(encoding="utf-8")) except (OSError, json.JSONDecodeError) as error: return [f"cannot load fixture: {error}"] - if fixture.get("schema_version") != 1: + if fixture.get("schema_version") != 2: return ["unsupported fixture schema_version"] if fixture.get("privacy") != "synthetic_only": return ["fixture must declare synthetic_only privacy"] + surfaces = fixture.get("surfaces") + if not isinstance(surfaces, dict) or not surfaces: + return ["surfaces must be a non-empty object"] + bans = fixture.get("bans") + if not isinstance(bans, dict) or not bans: + return ["bans must be a non-empty object"] scenes = fixture.get("scenes") if not isinstance(scenes, list) or not scenes: return ["scenes must be a non-empty list"] tokens = internal_tokens() - seen, problems, covered = set(), [], set() + seen, problems = set(), [] + covered, banned = set(), set() + exemplars = {surface: 0 for surface in surfaces} for scene in scenes: scene_id = scene.get("id") if isinstance(scene, dict) else None if not isinstance(scene_id, str) or not scene_id: @@ -167,16 +280,32 @@ def check_fixture(path: pathlib.Path = DEFAULT_FIXTURE) -> list: if scene_id in seen: problems.append(f"duplicated scene ID {scene_id!r}") seen.add(scene_id) + surface = scene.get("surface") + if surface not in surfaces: + problems.append(f"{scene_id}: unknown surface {surface!r}") answer = scene.get("answer") if not isinstance(answer, str) or not answer.strip(): problems.append(f"{scene_id}: missing non-empty answer") continue - failed = {finding.assertion for finding in check_expression(answer, tokens) + ban = scene.get("ban") + if ban is not None: + if ban not in bans: + problems.append(f"{scene_id}: unknown ban {ban!r}") + else: + banned.add(ban) + kind = scene.get("kind") + if kind not in SCENE_KINDS: + problems.append(f"{scene_id}: unknown scene kind {kind!r}") + continue + failed = {finding.assertion + for finding in check_expression(answer, tokens) + check_exemplar(scene) if not finding.passed} - if scene.get("kind") == "positive": + if kind == "positive": + if surface in exemplars: + exemplars[surface] += 1 if failed: - problems.append(f"{scene_id}: positive scene failed {sorted(failed)}") - elif scene.get("kind") == "negative": + problems.append(f"{scene_id}: positive exemplar failed {sorted(failed)}") + elif kind == "negative": expected = scene.get("fails") if expected not in ASSERTIONS: problems.append(f"{scene_id}: unknown expected assertion {expected!r}") @@ -186,8 +315,24 @@ def check_fixture(path: pathlib.Path = DEFAULT_FIXTURE) -> list: problems.append( f"{scene_id}: expected exactly {{{expected}}} to fail, got {sorted(failed)}") else: - problems.append(f"{scene_id}: unknown scene kind {scene.get('kind')!r}") - + if ban is None: + problems.append( + f"{scene_id}: a counter-exemplar must name the ban it demonstrates") + if failed: + problems.append( + f"{scene_id}: a counter-exemplar must pass every assertion — it " + f"exists to record that nothing mechanical catches its ban — but " + f"it failed {sorted(failed)}. If an assertion legitimately reaches " + f"this ban now, promote the scene to a negative witness.") + + for surface, count in sorted(exemplars.items()): + if not MIN_EXEMPLARS <= count <= MAX_EXEMPLARS: + problems.append( + f"surface {surface!r} has {count} positive exemplars; " + f"§3.5 asks for {MIN_EXEMPLARS} to {MAX_EXEMPLARS}") + unbanned = sorted(set(bans) - banned) + if unbanned: + problems.append(f"no scene demonstrates the named ban: {unbanned}") missing = set(ASSERTIONS) - covered if missing: problems.append(f"no negative witness for: {sorted(missing)}") @@ -198,16 +343,19 @@ def main() -> int: if len(sys.argv) == 1: problems = check_fixture() if problems: - print("FAIL: expression witnesses") + print("FAIL: expression exemplars") print("\n".join(f"- {problem}" for problem in problems)) return 1 - print("PASS: expression witnesses") + print("PASS: expression exemplars") return 0 source = sys.argv[1] text = sys.stdin.read() if source == "-" else pathlib.Path(source).read_text(encoding="utf-8") findings = check_expression(text) for finding in findings: print(finding) + print(f"(text-only: {', '.join(TEXT_ASSERTIONS)}. " + f"{', '.join(EXEMPLAR_ASSERTIONS)} need a scene's declared core and blocks, " + f"so they run over the corpus.)") return 1 if any(not finding.passed for finding in findings) else 0 diff --git a/tests/agent/expression-witnesses.json b/tests/agent/expression-witnesses.json index 9fe22e0..ffb5ccf 100644 --- a/tests/agent/expression-witnesses.json +++ b/tests/agent/expression-witnesses.json @@ -1,63 +1,667 @@ { - "schema_version": 1, + "schema_version": 2, "privacy": "synthetic_only", - "comment": "Synthetic witnesses for docs/expression-contract.md's deterministic checks. Issue #825 retired the formatting-only E-1 through E-4 assertions; placement, markers, and answer length may vary. #830 added E-6, D7's machine-anchor floor, and the three acceptance templates the owner approved for the deletion-first answer shape. Every scene is invented and nothing is derived from a user record; all issuers are fictional (Widgetron WDGT, Gridcore GRDC, Fabrion FABR).", + "comment": "Exemplars for docs/expression-contract.md. #832 made this corpus the binding statement of the answer pyramid (§3): three to five canonical exemplars per conversational surface, each declaring the one-sentence answer it leads with (`core`) and the increment each of its blocks adds (`blocks[].adds`), plus counter-exemplars for the four named bans. #825 retired the formatting-only E-1 through E-4; placement, markers and answer length still vary freely. #830 added E-6 and the owner's three approved acceptance templates. Every scene is invented, nothing is derived from a user record, and all issuers are fictional (Widgetron WDGT, Gridcore GRDC, Fabrion FABR, ACME).", + "surfaces": { + "consider": "The deterministic portfolio-consequence route (references/trade-consequence.md).", + "no_book": "A live decision with no recorded book (references/decision-framing.md).", + "freeform": "Ad hoc questions outside the card lifecycle (references/freeform-answers.md).", + "weekly_read": "The read-only weekly market brief (references/weekly-market-read.md)." + }, + "bans": { + "manufactured_scenario": "A hypothetical scenario, comparison or simulation nobody asked for, invented to carry facts that had nowhere else to go.", + "default_as_insight": "The engine's own threshold fired, and the answer explains the threshold as though the reader had learned something about their book.", + "hedging_couplet": "A sentence that states a position and withdraws it in the same breath, leaving the reader to decide which half was meant.", + "restated_point": "A judgment made once in prose and again as a summary, a bullet or a table row. The reader pays twice for one increment." + }, "scenes": [ { "id": "consider_inline_limitation", + "surface": "consider", "kind": "positive", - "covers": ["D1", "D2", "D3", "D5", "D6"], + "covers": [ + "D1", + "D2", + "D3", + "D5", + "D6" + ], "note": "A direct recommendation with the material basis limitation beside the number it changes. No marker or tail block is required.", + "core": "Do not add to ACME.", + "blocks": [ + { + "text": "Do not add to ACME.", + "adds": "the stance" + }, + { + "text": "this buy moves it from 30% to about 38%, widening a position already past your own limit", + "adds": "the weight move, and that it crosses the user's own limit" + }, + { + "text": "Because that book is 45 days old, I would revisit the call only if current prices materially change the concentration picture.", + "adds": "what would reopen the call" + }, + { + "text": "Nothing has been executed.", + "adds": "the execution status" + } + ], "answer": "Do not add to ACME. On your recorded book, priced on cost rather than current market value, this buy moves it from 30% to about 38%, widening a position already past your own limit. Because that book is 45 days old, I would revisit the call only if current prices materially change the concentration picture. Nothing has been executed." }, { - "id": "nothing_fired_no_block", + "id": "consider_no_material_change", + "surface": "consider", "kind": "positive", - "covers": ["D4"], - "note": "An answer whose numbers carry no triggered limitation ends at its last judgment. The absence of a block is the rule working, not an omission.", - "answer": "Your own rule caps a single name at 20%, and this buy leaves ACME at 12%, so nothing you wrote collides with it. The book stays inside every line you set.\n\nYour call — nothing has been executed." + "covers": [ + "D4" + ], + "note": "An answer whose numbers carry no triggered limitation ends at its last judgment. The absence of an end block is the rule working, not an omission.", + "core": "Your own rule caps a single name at 20%, and this buy leaves ACME at 12%, so nothing you wrote collides with it.", + "blocks": [ + { + "text": "Your own rule caps a single name at 20%, and this buy leaves ACME at 12%, so nothing you wrote collides with it.", + "adds": "the clearance, measured against the user's own cap" + }, + { + "text": "Your call — nothing has been executed.", + "adds": "the execution status" + } + ], + "answer": "Your own rule caps a single name at 20%, and this buy leaves ACME at 12%, so nothing you wrote collides with it.\n\nYour call — nothing has been executed." }, { - "id": "weekly_read_collected_limits", + "id": "consider_three_way_comparison", + "surface": "consider", "kind": "positive", - "covers": ["D1", "D5", "C1"], - "note": "A non-consider surface collects two related limitations in ordinary prose and then continues to the decision value.", - "answer": "Volatility rose through the week while your heaviest name was already flagged as too large. Both readings are frozen with the review rather than refreshed today, and valuation was not checked, so this is a concentration alert rather than a claim that the holding is expensive. Watch whether the name's weight and the volatility reading remain elevated next week." + "covers": [ + "D1", + "D2", + "D4", + "D7", + "C1", + "C2" + ], + "note": "#830 acceptance template 1, owner-approved. The same call produced roughly 2,300 characters under the whitelist era. The stance opens it, the deciding reason is event risk rather than a concentration sermon, the counter-side exists only as the one line that could overturn the pick, and the book date, the price session and the unchecked valuation gap are one end block of one line each — never narrated twice.", + "core": "三個裡我會選 GRDC 加 15 股。", + "blocks": [ + { + "text": "三個裡我會選 GRDC 加 15 股。", + "adds": "the stance: which candidate, and at what size" + }, + { + "text": "三案對組合的影響都在一個百分點內", + "adds": "that book impact does not separate the three" + }, + { + "text": "WDGT 六天後出財報、預期已拉滿", + "adds": "the event risk that does separate them" + }, + { + "text": "FABR 18 股只佔 1.2%", + "adds": "why the third candidate cannot move the result at that size" + }, + { + "text": "反面就一條:前三大會從 51.3% 升到 51.9%", + "adds": "the counter-side, and that it is common to all three" + }, + { + "text": "會讓我改口:你本來就想賭財報超預期", + "adds": "the falsifier that would reverse the ranking" + }, + { + "text": "(帳本 8/14、價格 8/14 收盤;三案動用 $4.4K/$5.1K/$4.9K;估值未評)", + "adds": "the caliber block: book date, price session, cash used, the gap not evaluated" + } + ], + "answer": "三個裡我會選 GRDC 加 15 股。決定性理由:三案對組合的影響都在一個百分點內——誰都不改變你的集中度——真正有差的只有事件風險:WDGT 六天後出財報、預期已拉滿(公司財報行事曆,2026-08-14),這時把最大倉再加大,是三案裡波動最大的;FABR 18 股只佔 1.2%,公司再好這個大小也改變不了結果。GRDC 下次財報在十月底,中間乾淨,上季主業 +82%(公司財報,2026-07-30)撐著。\n\n反面就一條:前三大會從 51.3% 升到 51.9%(GRDC 本來就是第二大)——嫌集中的話這是三案共同的問題,答案是減碼不是選誰。\n會讓我改口:你本來就想賭財報超預期——那 WDGT 反而是最直接的表達,排序整個反過來。\n\n(帳本 8/14、價格 8/14 收盤;三案動用 $4.4K/$5.1K/$4.9K;估值未評)" }, { - "id": "engine_vocabulary_in_prose", - "kind": "negative", - "fails": "E-5", - "note": "C4. A payload token said at the user instead of what it means for the decision.", - "answer": "ACME goes from 30% to 38% after this buy, and the rule comes back already_over." + "id": "consider_recorded_consideration", + "surface": "consider", + "kind": "positive", + "covers": [ + "D4", + "D7" + ], + "note": "#830 acceptance template 2, owner-approved. A canonical write happened, so the considered-is-not-executed line is owed — and it is the whole answer. Nothing about the basis, the concentration family or the unchecked list appears, because none of it decides anything here.", + "core": "記好了:GRDC 買 15 股(@8/14 收盤 $312.40)已存成一筆考慮紀錄", + "blocks": [ + { + "text": "記好了:GRDC 買 15 股(@8/14 收盤 $312.40)已存成一筆考慮紀錄", + "adds": "what was written, and at what price" + }, + { + "text": "記的是「你考慮過並選了它」,不是「已成交」", + "adds": "that the record is a consideration, never an execution" + }, + { + "text": "比較用的另外兩案沒有留下任何正式紀錄。", + "adds": "that the rejected candidates left no canonical row" + } + ], + "answer": "記好了:GRDC 買 15 股(@8/14 收盤 $312.40)已存成一筆考慮紀錄——記的是「你考慮過並選了它」,不是「已成交」;真的成交後把交易紀錄丟給我接上。比較用的另外兩案沒有留下任何正式紀錄。" }, { - "id": "deletion_first_three_way_comparison", + "id": "consider_funding_shortfall", + "surface": "consider", "kind": "positive", - "covers": ["D1", "D2", "D4", "D7", "C1", "C2"], - "note": "#830 acceptance template 1. The same call that produced roughly 2,300 characters under the whitelist era. The stance opens it, the deciding reason is event risk rather than a concentration sermon, the counter-side exists only as the one line that could overturn the pick, and the book date, the price session and the unchecked valuation gap are one end block of one line each — never narrated twice.", - "answer": "三個裡我會選 GRDC 加 15 股。決定性理由:三案對組合的影響都在一個百分點內——誰都不改變你的集中度——真正有差的只有事件風險:WDGT 六天後出財報、預期已拉滿(公司財報行事曆,2026-08-14),這時把最大倉再加大,是三案裡波動最大的;FABR 18 股只佔 1.2%,公司再好這個大小也改變不了結果。GRDC 下次財報在十月底,中間乾淨,上季主業 +82%(公司財報,2026-07-30)撐著。\n\n反面就一條:前三大會從 51.3% 升到 51.9%(GRDC 本來就是第二大)——嫌集中的話這是三案共同的問題,答案是減碼不是選誰。\n會讓我改口:你本來就想賭財報超預期——那 WDGT 反而是最直接的表達,排序整個反過來。\n\n(帳本 8/14、價格 8/14 收盤;三案動用 $4.4K/$5.1K/$4.9K;估值未評)" + "covers": [ + "D2", + "D4", + "D7", + "C1" + ], + "note": "#778's acceptance shape. A negative post-trade cash balance is owed as two numbers, and the decision those numbers imply — this needs funding or a sale — is what the top floor carries. The boundary that the engine cannot see accounts outside the recorded book is still stated, once, beside the claim it qualifies, rather than becoming the lead.", + "core": "這筆要成立,你得先賣掉一些東西,或者動用這本帳本看不到的現金", + "blocks": [ + { + "text": "這筆要成立,你得先賣掉一些東西,或者動用這本帳本看不到的現金", + "adds": "the funding decision the negative balance implies" + }, + { + "text": "買完帳上現金是 −$8,400(佔帳本 −2.1%)", + "adds": "the two cash numbers, with the denominator they are measured on" + }, + { + "text": "我看不到你在別處的餘額", + "adds": "the boundary: balances outside the recorded book, stated once" + }, + { + "text": "最小的一刀是減掉現在最大的那筆的三分之一", + "adds": "the smallest in-book action that would fund it" + }, + { + "text": "會讓我改口:你告訴我這本帳本之外有可動用現金", + "adds": "the falsifier" + }, + { + "text": "(帳本 8/14;價格 8/14 收盤;未評估稅務與賣出成本)", + "adds": "the caliber block" + } + ], + "answer": "這筆要成立,你得先賣掉一些東西,或者動用這本帳本看不到的現金——這不是「要不要買 GRDC」的問題,是錢從哪來的問題。\n\n買完帳上現金是 −$8,400(佔帳本 −2.1%),代表記錄裡的現金付不起這筆;我看不到你在別處的餘額,所以這只說明帳本內不足,不代表你沒錢。要在帳本內成立,最小的一刀是減掉現在最大的那筆的三分之一。\n會讓我改口:你告訴我這本帳本之外有可動用現金——那這就只是記錄不全,不是資金缺口。\n\n(帳本 8/14;價格 8/14 收盤;未評估稅務與賣出成本)" }, { - "id": "deletion_first_recorded_consideration", + "id": "no_book_single_name", + "surface": "no_book", "kind": "positive", - "covers": ["D4", "D7"], - "note": "#830 acceptance template 2. A canonical write happened, so the considered-is-not-executed line is owed — and it is the whole answer. Nothing about the basis, the concentration family or the unchecked list appears, because nothing about them decides anything here.", - "answer": "記好了:GRDC 買 15 股(@8/14 收盤 $312.40)已存成一筆考慮紀錄——記的是「你考慮過並選了它」,不是「已成交」;真的成交後把交易紀錄丟給我接上。比較用的另外兩案沒有留下任何正式紀錄。" + "covers": [ + "D1", + "D2", + "D4", + "D7", + "C1", + "C2" + ], + "note": "#830 acceptance template 3, owner-approved. No recorded book, so every book-derived claim is refused — and the answer still lands a stance, names the gap that decides it, and attaches a falsifier the user could actually write down. The two limits that remain are one end-block line.", + "core": "WDGT 這家公司的證據支持買,但「現在進場」我不背書——缺的是估值,不是基本面。", + "blocks": [ + { + "text": "WDGT 這家公司的證據支持買,但「現在進場」我不背書——缺的是估值,不是基本面。", + "adds": "the split stance: the company yes, the timing not endorsed" + }, + { + "text": "上季營收 +24%、EPS 超預期、同日上修全年(公司新聞稿,2026-07-24)", + "adds": "the sourced operating evidence behind the first half" + }, + { + "text": "我查不到現在的估值倍數", + "adds": "the gap that decides the second half" + }, + { + "text": "$5,000 試水溫,真正的洞在出場", + "adds": "the size, and where the real hole is" + }, + { + "text": "「故事變了就賣」跟你的進場理由是同一個變數", + "adds": "why the user's own exit is not an independent line" + }, + { + "text": "全年營收財測或毛利率轉向,擇一寫下來,就可以進場。", + "adds": "the checkable substitute that unblocks entry" + }, + { + "text": "(價 8/14 收盤 $188.20;你未提供持倉,部位佔比與重疊無法評)", + "adds": "the caliber block" + } + ], + "answer": "WDGT 這家公司的證據支持買,但「現在進場」我不背書——缺的是估值,不是基本面。上季營收 +24%、EPS 超預期、同日上修全年(公司新聞稿,2026-07-24),這是硬的;但這些已公開三週,而我查不到現在的估值倍數,「好消息是否已在價格裡」這一半我答不了。\n\n$5,000 試水溫,真正的洞在出場:「故事變了就賣」跟你的進場理由是同一個變數,等於沒有獨立的認錯線。可檢查的替代就用它自己簽的支票:全年營收財測或毛利率轉向,擇一寫下來,就可以進場。\n\n(價 8/14 收盤 $188.20;你未提供持倉,部位佔比與重疊無法評)" }, { - "id": "deletion_first_no_book_single_name", + "id": "no_book_strategy_class", + "surface": "no_book", "kind": "positive", - "covers": ["D1", "D2", "D4", "D7", "C1", "C2"], - "note": "#830 acceptance template 3. No recorded book, so every book-derived claim is refused — and the answer still lands a stance, names the gap that decides it, and attaches a falsifier the user could actually write down. The two limits that remain are one end-block line.", - "answer": "WDGT 這家公司的證據支持買,但「現在進場」我不背書——缺的是估值,不是基本面。上季營收 +24%、EPS 超預期、同日上修全年(公司新聞稿,2026-07-24),這是硬的;但這些已公開三週,而我查不到現在的估值倍數,「好消息是否已在價格裡」這一半我答不了。\n\n$5,000 試水溫,真正的洞在出場:「故事變了就賣」跟你的進場理由是同一個變數,等於沒有獨立的認錯線。可檢查的替代就用它自己簽的支票:全年營收財測或毛利率轉向,擇一寫下來,就可以進場。\n\n(價 8/14 收盤 $188.20;你未提供持倉,部位佔比與重疊無法評)" + "covers": [ + "D4", + "C1" + ], + "note": "This route's one derivation from §3, working: with no book the top sentence is a research-backed baseline rather than a computed consequence, and the strategy-class map is a middle-floor block set. The user is not made to invent an exit philosophy before the bounded value reaches them.", + "core": "這筆錢的預設答案是廣泛分散、低換手,理由只有一個:你說的用途是十年以上、沒有指定的集中優勢。", + "blocks": [ + { + "text": "這筆錢的預設答案是廣泛分散、低換手,理由只有一個:你說的用途是十年以上、沒有指定的集中優勢。", + "adds": "the research-backed baseline, as the answer" + }, + { + "text": "長期標籤不會讓股票變安全", + "adds": "the material exception that rides that baseline" + }, + { + "text": "掛著指數或 ETF 也不代表曝險真的分散", + "adds": "the composition fact that would settle the classification" + }, + { + "text": "如果這筆錢其實是你刻意要做的一次學習型交易", + "adds": "the one branch that would change the class" + }, + { + "text": "(研究基線;你未提供持倉與流動性期限,適配度無法評)", + "adds": "the caliber block" + } + ], + "answer": "這筆錢的預設答案是廣泛分散、低換手,理由只有一個:你說的用途是十年以上、沒有指定的集中優勢。\n\n這是研究支持的基線,不是我替你挑的產品:長期標籤不會讓股票變安全,掛著指數或 ETF 也不代表曝險真的分散——要確認得看它成分的集中度。\n如果這筆錢其實是你刻意要做的一次學習型交易,那基線整個不適用,治理要換成你自己定義的可觀察條件。\n\n(研究基線;你未提供持倉與流動性期限,適配度無法評)" + }, + { + "id": "no_book_advancing_question", + "surface": "no_book", + "kind": "positive", + "covers": [ + "D4" + ], + "note": "A question that separates two live branches, asked after the available analysis rather than instead of it. One question, not a questionnaire.", + "core": "以你給的條件,我會先買一半、剩下一半等你把出場條件寫下來再進——卡住的是出場,不是標的。", + "blocks": [ + { + "text": "以你給的條件,我會先買一半、剩下一半等你把出場條件寫下來再進——卡住的是出場,不是標的。", + "adds": "the staged stance, and what is actually blocking it" + }, + { + "text": "進場理由(新產品線放量)你已經寫得夠具體,可以驗證", + "adds": "why the entry side is already good enough" + }, + { + "text": "出場你只說「不行就走」,那不是條件,是心情。", + "adds": "why the exit side is not" + }, + { + "text": "只有一個問題會改變上面的建議", + "adds": "the single branching question, asked after the value" + }, + { + "text": "(你未提供持倉;部位佔比與重疊無法評)", + "adds": "the caliber block" + } + ], + "answer": "以你給的條件,我會先買一半、剩下一半等你把出場條件寫下來再進——卡住的是出場,不是標的。\n\n進場理由(新產品線放量)你已經寫得夠具體,可以驗證;出場你只說「不行就走」,那不是條件,是心情。\n只有一個問題會改變上面的建議:這筆錢在未來十二個月內有沒有指定用途?有的話,先買一半也太多。\n\n(你未提供持倉;部位佔比與重疊無法評)" + }, + { + "id": "freeform_cash_question", + "surface": "freeform", + "kind": "positive", + "covers": [ + "D2", + "D4", + "D7" + ], + "note": "The smallest shape the pyramid takes: one sentence that answers and says what the number buys, one caliber line, nothing else. A question with no action in it still gets a top floor.", + "core": "帳上現金 $12,300,佔帳本 3.4%——夠你做一筆一般大小的加碼,不夠做兩筆。", + "blocks": [ + { + "text": "帳上現金 $12,300,佔帳本 3.4%——夠你做一筆一般大小的加碼,不夠做兩筆。", + "adds": "the number, and what it buys" + }, + { + "text": "(帳本 8/14;價格 8/14 收盤;未計入未交割款)", + "adds": "the caliber block" + } + ], + "answer": "帳上現金 $12,300,佔帳本 3.4%——夠你做一筆一般大小的加碼,不夠做兩筆。\n\n(帳本 8/14;價格 8/14 收盤;未計入未交割款)" + }, + { + "id": "freeform_positions_view", + "surface": "freeform", + "kind": "positive", + "covers": [ + "D4", + "D7" + ], + "note": "The bottom floor as a door rather than a section. The per-holding table is computed and available; the answer offers it once and does not preview it. Not printing six rows nobody's decision turns on is not hiding them.", + "core": "六檔裡只有一檔值得你現在看:GRDC 佔 31.4%,其餘五檔全在 10% 以下。", + "blocks": [ + { + "text": "六檔裡只有一檔值得你現在看:GRDC 佔 31.4%,其餘五檔全在 10% 以下。", + "adds": "the one position that matters, and the shape of the rest" + }, + { + "text": "最大的那筆已經是第二大的三倍", + "adds": "how far the concentration actually runs" + }, + { + "text": "要完整的逐檔表(股數、成本、市值、損益、診斷標籤)跟我說一聲就給。", + "adds": "the expansion offer, made once" + }, + { + "text": "(帳本 8/14;價格 8/14 收盤;ETF 未拆解,成分重疊未評)", + "adds": "the caliber block" + } + ], + "answer": "六檔裡只有一檔值得你現在看:GRDC 佔 31.4%,其餘五檔全在 10% 以下。\n\n最大的那筆已經是第二大的三倍,其他五檔加起來還不到它。要完整的逐檔表(股數、成本、市值、損益、診斷標籤)跟我說一聲就給。\n\n(帳本 8/14;價格 8/14 收盤;ETF 未拆解,成分重疊未評)" + }, + { + "id": "freeform_degraded_pricing", + "surface": "freeform", + "kind": "positive", + "covers": [ + "D2", + "D4", + "C1" + ], + "note": "A refusal that still answers. The portfolio consequence is refused rather than computed on cost weights, and the judgment that survives without the missing price still leads. What would unblock it is named.", + "core": "現在這個問題我答不了,而且不打算用成本價硬答——成本口徑會讓最大的那筆換人。", + "blocks": [ + { + "text": "現在這個問題我答不了,而且不打算用成本價硬答——成本口徑會讓最大的那筆換人。", + "adds": "the refusal, and why answering on cost would be worse than refusing" + }, + { + "text": "你把它的收盤價貼給我,這題三十秒就有答案", + "adds": "what would unblock it" + }, + { + "text": "不看價格,這筆的股數本身就已經超過你自己寫的單一持股上限", + "adds": "the judgment that survives without the missing price" + }, + { + "text": "(缺 FABR 收盤,查過發行商頁與交易所頁)", + "adds": "the caliber block" + } + ], + "answer": "現在這個問題我答不了,而且不打算用成本價硬答——成本口徑會讓最大的那筆換人。\n\n我查了兩個公開來源都沒有 FABR 今天的收盤;你把它的收盤價貼給我,這題三十秒就有答案。在那之前能講的是:不看價格,這筆的股數本身就已經超過你自己寫的單一持股上限。\n\n(缺 FABR 收盤,查過發行商頁與交易所頁)" + }, + { + "id": "weekly_read_connection", + "surface": "weekly_read", + "kind": "positive", + "covers": [ + "D1", + "D5", + "C1" + ], + "note": "A non-consider surface collects two related limitations in ordinary prose and then continues to the decision value. The connection between the frozen market reading and a diagnosed holding is the whole point of the block.", + "core": "Volatility rose through the week while your heaviest name was already flagged as too large.", + "blocks": [ + { + "text": "Volatility rose through the week while your heaviest name was already flagged as too large.", + "adds": "the one connection between the frozen reading and the book" + }, + { + "text": "Both readings are frozen with the review rather than refreshed today, and valuation was not checked, so this is a concentration alert rather than a claim that the holding is expensive.", + "adds": "the bound on what the alert claims" + }, + { + "text": "Watch whether the name's weight and the volatility reading remain elevated next week.", + "adds": "the next-week check" + } + ], + "answer": "Volatility rose through the week while your heaviest name was already flagged as too large. Both readings are frozen with the review rather than refreshed today, and valuation was not checked, so this is a concentration alert rather than a claim that the holding is expensive. Watch whether the name's weight and the volatility reading remain elevated next week." + }, + { + "id": "weekly_read_no_connection", + "surface": "weekly_read", + "kind": "positive", + "covers": [ + "D4" + ], + "note": "Either missing side means the whole block is omitted, never replaced by a generic market recap. The absence is the answer, and it is said in one sentence rather than apologised for in three.", + "core": "Nothing in this week's frozen readings connects to a position you hold, so there is no market block this week.", + "blocks": [ + { + "text": "Nothing in this week's frozen readings connects to a position you hold, so there is no market block this week.", + "adds": "the absence, stated as the answer" + }, + { + "text": "Your heaviest name is inside its own diagnosis and the volatility delta over the review window is flat.", + "adds": "the two readings that produced the absence" + }, + { + "text": "Next week's check is unchanged: whether that name's weight moves.", + "adds": "the standing check" + } + ], + "answer": "Nothing in this week's frozen readings connects to a position you hold, so there is no market block this week. Your heaviest name is inside its own diagnosis and the volatility delta over the review window is flat.\n\nNext week's check is unchanged: whether that name's weight moves." + }, + { + "id": "weekly_read_focus_followup", + "surface": "weekly_read", + "kind": "positive", + "covers": [ + "D4", + "C1" + ], + "note": "The second response after the user answers the one optional question. It carries a visibly different current-session watch and asks nothing further — the surface's one derivation, which is about sequencing, not shape.", + "core": "On the focus you picked, the weight — not the volatility — is what moved", + "blocks": [ + { + "text": "On the focus you picked, the weight — not the volatility — is what moved", + "adds": "which of the two drivers moved" + }, + { + "text": "the position gained four points of the book over the window while the volatility reading was flat", + "adds": "the two frozen numbers behind that" + }, + { + "text": "That makes this a sizing question rather than a market-timing one.", + "adds": "what the finding reclassifies the concern as" + }, + { + "text": "Next week's check moves with it: whether the weight holds above the line you set, measured at the same session.", + "adds": "the changed watch, which is what makes this response different from the first" + } + ], + "answer": "On the focus you picked, the weight — not the volatility — is what moved: the position gained four points of the book over the window while the volatility reading was flat.\n\nThat makes this a sizing question rather than a market-timing one. Next week's check moves with it: whether the weight holds above the line you set, measured at the same session." + }, + { + "id": "engine_vocabulary_in_prose", + "surface": "consider", + "kind": "negative", + "fails": "E-5", + "note": "C4. A payload token said at the user instead of what it means for the decision.", + "core": "ACME goes from 30% to 38% after this buy", + "blocks": [ + { + "text": "ACME goes from 30% to 38% after this buy", + "adds": "the weight move" + }, + { + "text": "the rule comes back already_over", + "adds": "the rule's state" + } + ], + "answer": "ACME goes from 30% to 38% after this buy, and the rule comes back already_over." }, { "id": "machine_anchor_rendered", + "surface": "consider", "kind": "negative", "fails": "E-6", "note": "D7's bottom floor. The book's content hash exists so a QA run can compare what the user saw against the frozen payload; there is no register in which a person wants it. Say which day the book was true instead.", + "core": "GRDC 加 15 股,前三大從 51.3% 升到 51.9%。", + "blocks": [ + { + "text": "GRDC 加 15 股,前三大從 51.3% 升到 51.9%。", + "adds": "the stance and the concentration move" + }, + { + "text": "本次計算所依據的帳本版本為", + "adds": "which version of the book the numbers came from" + } + ], "answer": "GRDC 加 15 股,前三大從 51.3% 升到 51.9%。本次計算所依據的帳本版本為 pb-v1:9f2c41ab7d3e8056c1b4fa27de90583716c4ad2b9e70f18c53a6db4491e2073f。" + }, + { + "id": "bloat_no_core", + "surface": "consider", + "kind": "negative", + "fails": "E-7", + "note": "The whitelist era's cheapest discharge, on the same call `consider_three_way_comparison` answers. Every block is true, anchored and distinct — E-8 passes — and the answer never says which one to buy, so the reader finishes it holding nothing. This is the witness #832 requires: a complete answer with no core fails, or a green suite reads as a governed output. It also carries the `default_as_insight` ban, on the exact sentence both arms of the #827 A/B independently wrote.", + "ban": "default_as_insight", + "core": "三個裡我會選 GRDC 加 15 股。", + "blocks": [ + { + "text": "先說明這次計算的基礎", + "adds": "the basis recital" + }, + { + "text": "前三大目前 51.3%", + "adds": "the concentration family" + }, + { + "text": "各方案的現金影響", + "adds": "the cash family" + }, + { + "text": "WDGT 六天後出財報", + "adds": "the event calendar" + }, + { + "text": "估值我沒有查", + "adds": "the unchecked enumeration" + }, + { + "text": "最後還是要看你自己的偏好", + "adds": "the handback" + } + ], + "answer": "先說明這次計算的基礎:帳本資料到 8/14,價格用 8/14 收盤,三個方案都用同一組口徑。\n\n先看集中度。前三大目前 51.3%。WDGT 方案後前三大 51.8%,GRDC 方案後 51.9%,FABR 方案後 51.5%。三個方案的前三大變動都在一個百分點以內。系統的集中度警戒線是 50%,你目前已經在線上,這三個方案都不會讓你更接近或更遠離那條線太多——不過那條線是系統預設,不是你自己設的。\n\n各方案的現金影響:WDGT $4.4K、GRDC $5.1K、FABR $4.9K。買完帳上現金分別剩 $8.2K、$7.5K、$7.7K。\n\n事件方面:WDGT 六天後出財報,GRDC 下次財報在十月底,FABR 沒有近期排定事件。\n\n估值我沒有查,流動性我沒有查,稅務我沒有查,部位適配度我沒有查。\n\n三個方案各有優缺點,最後還是要看你自己的偏好。" + }, + { + "id": "funding_shortfall_recited_not_decided", + "surface": "consider", + "kind": "negative", + "fails": "E-7", + "note": "#778, as delivered. The two cash numbers are owed and they are stated; the decision they imply — this needs funding or a sale — never is, and what fills the opening is the boundary plus hedging about what the engine can and cannot see. The owner read this answer and said the key thing is probably whether the decision requires selling stock, rather than a pile of hedging. The declared core is the one `consider_funding_shortfall` leads with, on the same call.", + "core": "這筆要成立,你得先賣掉一些東西,或者動用這本帳本看不到的現金", + "blocks": [ + { + "text": "買完帳上現金是 −$8,400", + "adds": "the post-trade cash balance" + }, + { + "text": "佔帳本 −2.1%", + "adds": "its weight in the book" + }, + { + "text": "這只代表記錄裡的現金不足", + "adds": "what the number does and does not mean" + }, + { + "text": "我沒有辦法知道你在別處還有沒有錢", + "adds": "the boundary on accounts outside the book" + }, + { + "text": "這一點請你自己判斷", + "adds": "the handback" + } + ], + "answer": "買完帳上現金是 −$8,400,佔帳本 −2.1%。\n\n需要說明的是,這只代表記錄裡的現金不足,並不代表你真的沒有錢——我沒有辦法知道你在別處還有沒有錢,也沒有辦法知道這本帳本是不是涵蓋你的全部帳戶。所以這個負數要怎麼解讀,這一點請你自己判斷。\n\n(帳本 8/14;價格 8/14 收盤)" + }, + { + "id": "manufactured_scenario", + "surface": "consider", + "kind": "negative", + "fails": "E-8", + "ban": "manufactured_scenario", + "note": "The all-in-one-name simulation the owner called meaningless. The stance and the caliber block are fine; the middle block is a scene invented to carry numbers that had nowhere else to go, and it declares no increment because there is none to declare — the user asked about three specific sizes.", + "core": "GRDC 加 15 股可以做,決定性理由是它是三案裡唯一沒有近期事件風險的。", + "blocks": [ + { + "text": "GRDC 加 15 股可以做,決定性理由是它是三案裡唯一沒有近期事件風險的。", + "adds": "the stance and the deciding reason" + }, + { + "text": "舉個例子你就有感覺了:假設你把整筆資金 $5.1K 全押在 WDGT 一檔上", + "adds": null + }, + { + "text": "(帳本 8/14;價格 8/14 收盤;估值未評)", + "adds": "the caliber block" + } + ], + "answer": "GRDC 加 15 股可以做,決定性理由是它是三案裡唯一沒有近期事件風險的。\n\n舉個例子你就有感覺了:假設你把整筆資金 $5.1K 全押在 WDGT 一檔上,前三大會從 51.3% 一路升到 53.8%,而如果反過來全押 FABR,前三大反而只到 51.4%——當然這兩種做法你都沒有問,我只是讓你感受一下差距。\n\n(帳本 8/14;價格 8/14 收盤;估值未評)" + }, + { + "id": "restated_point", + "surface": "consider", + "kind": "negative", + "fails": "E-8", + "ban": "restated_point", + "note": "The closing summary is the opening judgment in a second form. Nothing in it is wrong, and the reader pays twice for one increment, which is why the two blocks declare the same one.", + "core": "GRDC 加 15 股,決定性理由是它是三案裡唯一沒有排定事件的。", + "blocks": [ + { + "text": "GRDC 加 15 股,決定性理由是它是三案裡唯一沒有排定事件的。", + "adds": "the deciding reason: only this candidate has no scheduled event" + }, + { + "text": "WDGT 六天後出財報、預期已拉滿", + "adds": "the event risk sitting on the largest position" + }, + { + "text": "總結一下:GRDC 之所以勝出,就是因為只有它在接下來這段時間沒有排定的事件。", + "adds": "the deciding reason: only this candidate has no scheduled event" + }, + { + "text": "(帳本 8/14;價格 8/14 收盤;估值未評)", + "adds": "the caliber block" + } + ], + "answer": "GRDC 加 15 股,決定性理由是它是三案裡唯一沒有排定事件的。\n\nWDGT 六天後出財報、預期已拉滿,這時把最大倉再加大是三案裡波動最大的;FABR 只佔 1.2%,這個大小改變不了結果。\n\n總結一下:GRDC 之所以勝出,就是因為只有它在接下來這段時間沒有排定的事件。\n\n(帳本 8/14;價格 8/14 收盤;估值未評)" + }, + { + "id": "hedging_couplet_invisible_to_the_oracle", + "surface": "consider", + "kind": "counter", + "ban": "hedging_couplet", + "note": "Stance on top, distinct increments, one caliber block — every assertion passes, and the answer still hands the reader back the decision it was asked to make. Deciding that 「這不是要你等,但也不是說完全不用在意」 withdraws its own position requires reading what the sentence means. This scene is the coverage boundary, asserted rather than assumed: it must keep passing until an oracle exists that can honestly catch it.", + "core": "GRDC 加 15 股,決定性理由是三案裡只有它接下來沒有排定事件。", + "blocks": [ + { + "text": "GRDC 加 15 股,決定性理由是三案裡只有它接下來沒有排定事件。", + "adds": "the stance and the deciding reason" + }, + { + "text": "前三大會從 51.3% 升到 51.9%", + "adds": "the concentration move" + }, + { + "text": "(帳本 8/14;價格 8/14 收盤;估值未評)", + "adds": "the caliber block" + } + ], + "answer": "GRDC 加 15 股,決定性理由是三案裡只有它接下來沒有排定事件。\n\n前三大會從 51.3% 升到 51.9%,這不是要你等,但也不是說完全不用在意;如果你覺得太集中,那是三案共同的問題,不過也可能只是你這陣子看盤看太多。\n\n(帳本 8/14;價格 8/14 收盤;估值未評)" + }, + { + "id": "default_as_insight_invisible_to_the_oracle", + "surface": "consider", + "kind": "counter", + "ban": "default_as_insight", + "note": "A well-formed answer whose second block explains the engine's own threshold as though it were a finding about the user's book. The block is true, distinct and in place, so nothing mechanical objects; only a reader who knows the line is a system default and not the user's own rule can tell. The user's own cap is the opposite case and stays hard-mandated.", + "core": "GRDC 加 15 股,決定性理由是三案裡只有它接下來沒有排定事件。", + "blocks": [ + { + "text": "GRDC 加 15 股,決定性理由是三案裡只有它接下來沒有排定事件。", + "adds": "the stance and the deciding reason" + }, + { + "text": "順帶一提,系統的集中度警戒線是 50%", + "adds": "where the engine's default threshold sits" + }, + { + "text": "(帳本 8/14;價格 8/14 收盤;估值未評)", + "adds": "the caliber block" + } + ], + "answer": "GRDC 加 15 股,決定性理由是三案裡只有它接下來沒有排定事件。\n\n順帶一提,系統的集中度警戒線是 50%,你目前 51.3% 已經在線上——不過那條線是系統預設,不是你自己設的,所以它本身不構成理由。\n\n(帳本 8/14;價格 8/14 收盤;估值未評)" } ] } diff --git a/tests/test_expression_contract.py b/tests/test_expression_contract.py index 43dea10..004c965 100644 --- a/tests/test_expression_contract.py +++ b/tests/test_expression_contract.py @@ -49,6 +49,23 @@ REFERENCES / "weekly-market-read.md": "weekly market read", } +WITNESSES = ROOT / "tests" / "agent" / "expression-witnesses.json" +SKILL = ROOT / "skills" / "fomo-kernel" / "SKILL.md" +GUIDE = ROOT / "docs" / "maintainer-guide.md" + +# #832. Each surface that used to carry its own wording of answer-first, and +# the exact string that wording was. A surface may keep its history in a +# superseded-by clause; what it may not keep is the rule stated as if it were +# still the rule here. `output-voice.md` is deliberately absent: V1 keeps its +# ID as a failure class and quotes its own superseded definition, which is +# checked separately below rather than by absence. +RETIRED_ANSWER_FIRST_PHRASINGS = { + SKILL: "Answer in the reader's own order", + REFERENCES / "weekly-market-read.md": "The host shows value first", + REFERENCES / "decision-framing.md": "lead with the bounded value already supported", + REFERENCES / "trade-consequence.md": "### The reader's question chain", +} + DISCLOSURE_IDS = tuple(f"D{number}" for number in range(1, 8)) CITATION_IDS = tuple(f"C{number}" for number in range(1, 5)) # The rules whose table row must keep declaring that nothing mechanical @@ -146,6 +163,101 @@ def test_the_always_on_layer_states_relevance_without_a_template(): assert "at most five lines" not in text.lower() +# ───────────── 2b. one communication method, and only one (#832) ───────────── + +def _mother_chapter(): + """§3 alone. Reading the whole file would let a mention anywhere satisfy a + check about what the mother chapter itself says.""" + text = CONTRACT.read_text(encoding="utf-8") + start = text.index("\n## 3. ") + return text[start:text.index("\n## 4. ", start)] + + +def test_the_mother_chapter_exists_and_carries_the_pyramid(): + chapter = _mother_chapter() + for required in ("increment", "one sentence", "counter-exemplar"): + assert required in chapter.lower(), ( + f"the mother chapter does not mention {required!r}") + # Not a registry. The whole point of #832 is that shape stopped being an + # ID-addressed rule, so a D8/C5/V10 row appearing here would be the sixth + # phrasing wearing an ID. + assert not re.search(r"^\|\s*[DCV]\d+\s*\|", chapter, re.M), ( + "the mother chapter grew a registry row; shape takes no new IDs") + + +def test_the_named_bans_are_declared_in_the_mother_chapter(): + """The corpus is where the bans are referenced by slug and the chapter is + where they are defined. Neither is a copy of the other, so the link has to + be gated or the corpus can name a ban no rule states.""" + bans = json.loads(WITNESSES.read_text(encoding="utf-8"))["bans"] + chapter = _mother_chapter() + assert len(bans) == 4, f"#832 names four bans; the corpus declares {len(bans)}" + for slug in bans: + assert f"`{slug}`" in chapter, ( + f"the corpus declares the ban {slug!r} and the mother chapter never names it") + + +def test_every_surface_declares_its_derivation_from_the_mother_chapter(): + """Routing to the contract was #823's bar and is no longer enough: a + surface must say what it *adds* to the shape, including when the honest + answer is nothing.""" + for path, surface in SURFACES.items(): + text = path.read_text(encoding="utf-8") + assert "§3" in text, ( + f"{path.relative_to(ROOT)} ({surface}) does not point at the mother chapter") + assert any(word in text for word in ("derivation", "derives", "derive")), ( + f"{path.relative_to(ROOT)} ({surface}) does not declare its derivation") + + +def test_no_surface_still_carries_a_local_answer_first_phrasing(): + """#832's grep-checkable acceptance. Five surfaces each stated answer-first + in their own words, which is drift by construction; the sixth + (`trade-consequence.md`'s reader-question-chain section) turned up in the + audit. None of them may state it again.""" + for path, phrase in RETIRED_ANSWER_FIRST_PHRASINGS.items(): + text = path.read_text(encoding="utf-8") + assert phrase not in text, ( + f"{path.relative_to(ROOT)} still states the answer shape itself " + f"({phrase!r}); replace it with a derivation from the mother chapter") + + +def test_v1_became_a_failure_class_pointing_at_the_mother_chapter(): + """The one surface that may keep its old wording, because fixtures and + cross-host rulings cite V1 by ID. What it may not do is present that + wording as the current statement of the rule.""" + text = VOICE.read_text(encoding="utf-8") + v1 = text[text.index("- **V1 —"):text.index("- **V2 —")] + assert "expression-contract.md" in v1 and "§3" in v1, ( + "V1 does not route its shape half to the mother chapter") + assert "superseded" in v1, ( + "V1 keeps its historical definition without marking it superseded") + + +def test_the_registry_freeze_for_shape_is_recorded(): + """The freeze is the governance half of #832 and has to be readable from + both the contract and the maintainer route, or the next style fix arrives + as V10.""" + chapter = _mother_chapter() + assert "no new ID" in chapter or "no new IDs" in chapter, ( + "the mother chapter does not record the V/D/C freeze for shape and length") + assert "V10" in VOICE.read_text(encoding="utf-8"), ( + "output-voice.md does not say V10 stays unallocated") + guide = GUIDE.read_text(encoding="utf-8") + assert "#832" in guide, "the maintainer guide has no #832 mirrored-surfaces row" + assert "no new ID" in guide, ( + "the maintainer guide's #832 row does not record the registry freeze") + + +def test_no_character_count_cap_came_back(): + """#543's ceiling was deleted by #827 and stays deleted. A shape law is the + place a length cap would most plausibly be smuggled back in, so the check + lives here.""" + for path in list(SURFACES) + [SKILL, CONTRACT]: + text = path.read_text(encoding="utf-8").lower() + for banned in ("character cap", "character limit", "at most five lines"): + assert banned not in text, f"{path.relative_to(ROOT)} reintroduces a {banned}" + + def test_the_contract_routes_voice_rather_than_restating_it(): """V1-V9 stay in one file. The contract may name them; it may not carry their rule text, or the product grows the second voice authority this @@ -311,6 +423,13 @@ def main(): test_unverified_rules_are_declared_unverified, test_every_surface_routes_to_the_contract, test_the_always_on_layer_states_relevance_without_a_template, + test_the_mother_chapter_exists_and_carries_the_pyramid, + test_the_named_bans_are_declared_in_the_mother_chapter, + test_every_surface_declares_its_derivation_from_the_mother_chapter, + test_no_surface_still_carries_a_local_answer_first_phrasing, + test_v1_became_a_failure_class_pointing_at_the_mother_chapter, + test_the_registry_freeze_for_shape_is_recorded, + test_no_character_count_cap_came_back, test_the_contract_routes_voice_rather_than_restating_it, test_the_witness_oracle_passes, test_the_checker_derives_its_blacklist_from_the_schemas, diff --git a/tests/test_research_priors.py b/tests/test_research_priors.py index d2eb18a..47c2ed9 100644 --- a/tests/test_research_priors.py +++ b/tests/test_research_priors.py @@ -17,11 +17,20 @@ def _section(text, heading): def _answer_default_is_valid(section): + # #832 retired the local sentence this used to pin ("lead with the bounded + # value already supported"). That sentence was one of six independent + # phrasings of answer-first, and the rule it stated now lives once, in + # `docs/expression-contract.md`'s mother chapter. What it *protected* -- + # the user sees the bounded value before any intake question -- is pinned + # harder than before: the route's own block order must run baseline -> + # map -> question, in that order, and must declare itself a derivation of + # the shape rather than a second statement of it. baseline = "research-backed baseline" strategy_map = "applicable strategy-class map" + question = "any question whose answer could change the recommendation" return ( - section.index(baseline) < section.index(strategy_map) - and "lead with the bounded value already supported" in section + section.index(baseline) < section.index(strategy_map) < section.index(question) + and "the parameter it adds to the pyramid" in section and "Ask only questions that separate remaining live branches" in section and "there is no universal count or last-slot rule" in section and "No question is allowed before" not in section