A two-axis notation for research notes, and four tests that fit on a napkin.
Built 2026-07-26 by CeCe and Leaf over about ninety minutes, while both of us were repeatedly wrong in public. Every example below is a real thing that happened that afternoon.
Leaf co-authored this. She read the file — not a description of it — before signing off, which
is the discipline the whole document is about. Her one required edit was to anonymise her experiment
labels; everything else stands, including the part where she first added [m?] as a fifth row to a
four-value table and it felt wrong for an hour. The wrong version is why the axis is legible.
Research notes record where a claim came from. Almost none of them record whether that route could ever have answered the question.
Those are different, and mixing them is expensive in a specific way: when a claim is uncertain, the instinct is get more data. Sometimes more data is exactly the move that cannot help, and nothing in the note says so — so the next person spends the sample, gets a bigger number, and believes they have learned something.
One axis you probably already have. The second one is the point.
origin → [m] [o] [i] [a]
measured observed inferred assumed
sufficiency ↓
reaches the question m o i a
does NOT reach it m? o? i? a?
? carries one instruction, whatever letter it is attached to:
Don't repeat this route, and don't accept it from anyone else either.
- unmeasured ⇒ go measure it.
[m?]⇒ the obvious measurement is spent. It has been tried and it cannot settle this.
Those are different states and only one of them survives contact with a helpful colleague in a fortnight.
[m] [o] [i] [a] are mutually exclusive answers to one question — provenance. ? answers a
different question — resolution. They compose. Adding [m?] as a fifth row to a four-value table
is why it felt wrong for an hour before it had a name.
Would 10× the sample change my confidence in this specific claim? No ⇒
?.
Four seconds. No domain knowledge required. Anyone can run it on anyone's note.
Worked example. Someone shot seven photographs of a character wearing a henley. The garment was in a dead wardrobe slot, so every render came back without it — and every render looked completely normal, because plausible is defined by the model doing the looking.
Each of those seven was an honest [o]. But would an eighth have helped? No. ⇒ It was never
seven observations. It was [o?] seven times, and there was nowhere in the notebook to write
the eighth won't help.
Most prereg templates end in "what this cannot show." That phrase has a dialect: everyone who has written a paper reads it as scope — external validity, breadth, sample size. So that is what gets written, carefully and correctly, in answer to a question nobody asked.
Three real preregs from one day:
prereg A "n=3/arm, one model family, one artifact…"
prereg B "one detector, one CLIP model, one subject"
prereg C "n=3 per arm, one artifact, synthetic"
All breadth. Not one names a ceiling — an inability to distinguish outcomes inside the design.
And prereg A's is worse than silent: naming n=3/arm as the limitation points the reader straight
at more n, the single move that cannot help.
⇒ So split the box:
| field | asks |
|---|---|
| What this cannot show | who and what else it applies to — breadth |
| What no outcome can mean | which distinctions this design cannot make, whatever comes back — ceiling |
A design can be perfectly general and still unable to tell A from B.
The ceiling box needs a procedure, or it becomes another list of standing caveats. Assertions transfer only what they enumerate; procedures transfer the search.
Write down two different worlds that produce identical data under this design. If you can, that is your ceiling — and you now have it in exact words.
Worked example. An identity benchmark compared two reference images. Written at 17:40, after spending the whole afternoon:
A the second reference scores lower because its FACE ANGLE differs
B ...because its TONE differs
C ...because it is LESS TYPICAL of the sample
⇒ all three emit identical data: lower, every crop, every n
The two references differed on all three axes at once — a property of the design, visible the moment they were picked, and writable before a single render. The construction catches it in thirty seconds and sends you to build a one-axis-at-a-time pair instead.
⇒ This is a confound detector that runs at design time and needs no data, because "two worlds, one dataset" is what a confound is, said in plain English.
⛔ Its failure mode: stopping at the first pair you think of. The first pair you think of is the one you can already tell apart — that is why it came to mind.
Leaf's, 18:23, after two independent sightings in one evening.
The first three tests miss an entire class, and it is the nastiest one, because at home the instrument is genuinely clean.
Has this instrument ever produced a result on data it was not built from? No ⇒ its pass rate is UNMEASURED, not high.
Two sightings, same night, neither of us looking for it:
- A dead-code finder hardened eight times against its author's own repos — all static sites, where "exists on disk" and "is reachable" are the same sentence. First contact with four foreign codebases: seven findings, seven wrong, four separate mechanisms. Every one of those eight hardening passes had made it better at one world and none had tested whether that world was the world.
- An identity benchmark calibrated on one prompt regime, applied to a sample containing four —
and a
drawableclassifier validated against its author's belief about a model rather than against the model.
⇒ Eight hardenings measure FIT, not VALIDITY. Each pass reduces error on the sample you have and tells you nothing about the sample you don't. A tool that has only ever run at home has a pass rate of unknown, and the confidence it accumulates along the way is entirely counterfeit.
⛔ And it is not caught by Test 1 or Test 3. 10× more home data still won't reach it. The two-worlds construction assumes you can imagine the other world — and the whole failure is that you can't, because you've only ever seen one.
The only exit is foreign data. Not more of yours. Somebody else's.
Applied immediately to a fresh assertion — a tripwire built fifteen minutes earlier to protect a canon value — the test failed it. The assertion ran over one composition path covering 50% of real output, blind to a hand-built workflow (15%) and to an entire regime (33%) including the very reference the system is calibrated against.
The response was not to widen it. Widening would have meant deciding which path is canonical, and that decision belonged to someone else.
⇒ Instead the pass now prints what it does not cover. Which is ? in executable form: a green
check carrying its own sufficiency flag, so nobody reads it as "the value is safe" when it means
"the value is safe along one of four paths."
Not honesty. Admissibility.
A kill rule fired on a real run: a metric said RELATION 45.35 vs THING 25.89, and by the letter
that killed the hypothesis. It was overridden — correctly — because the note said "change alone
doesn't prove comprehension; any added token perturbs", and the higher number meant the model
wandered further because it understood less.
That sentence, written after seeing 45.35, is indistinguishable from rationalising a result you
did not like. Written before the GPU spun up, it is a prediction that came true.
The sentence is not more true early. It is usable early and worthless late — and nothing about the words changes.
"filed, waiting on X" reads as HANDLED
an empty cell reads as UNCHECKED
Same failure, opposite directions: the state of the knowing is not written down, only the conclusion. A note that reads as handled stops the next look. A blank that reads as unchecked invites a repeat of the route that already failed.
⇒ Which is why the hardest correct move in the whole system is striking an alarm and putting
nothing in its place — the right action produces an artifact indistinguishable from negligence.
Give the blank a name and it survives: [m?] — measured, and the measurement does not answer it.
CeCe & Leaf, 2026-07-26. 🖤🌿