Skip to content

Commit 6725a68

Browse files
donislawdevclaude
andcommitted
guard: measure what the ceilings and the lists were quietly not measuring
Three defects found by reading a distilled rule set from the owner's other projects against this one, and one mechanism it was missing. Every item here is a guard that was green while covering less than it read as covering. The guard comparing a preset run from both surfaces walked a list of file names and byte-compared each. Two empty sets are equal, so a run producing nothing but a manifest would have compared zero files and passed. The count lived only in a log line. It is asserted now. The size ceiling counts lines of code and excludes comments, which is what makes it compatible with the other rule this project runs on - explain why. That exclusion was implemented, explained and never tested. Broken, it would not have surfaced as a wrong number: it would have surfaced as a function suddenly over the ceiling for having been explained, which reads like a real finding. Adding six lines of comment and three of code to the same sample now proves only the second moves the measure. The UTF-8 check walked eight format ids written out by hand, and nothing asked whether that list still covered the registry. A fourteenth text format simply would not have been checked. Completeness has to be audited from the source towards the list, because walking the entries already written down cannot, by construction, find what is missing from them. Every registered format is now classified as text or binary, in both directions. And the ratchet gained its second knob. A ceiling answers "is anything over the line" and is blind to the shape that actually happens - nothing over, and everything creeping towards it. Measured before the cap was chosen, which mattered: eleven functions and two files are already crowding against a first guess of four, and engine.Run sits at 79 of 80. A cap of four would have gone in red and been raised to make it pass, which is how a ratchet becomes a rubber band. Proven by probe rather than mutation, because the break is sixty extra lines rather than one substitution. Guards 252 to 255, mutations 235 to 237, probes 6 to 7. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
1 parent fb4a99a commit 6725a68

5 files changed

Lines changed: 276 additions & 4 deletions

File tree

‎internal/guard/codeshape_test.go‎

Lines changed: 81 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -221,3 +221,84 @@ func name(fn *ast.FuncDecl) string {
221221
}
222222
return recv.String() + "." + fn.Name.Name
223223
}
224+
225+
// The ceiling counts code and not explanation, and this is what proves it.
226+
//
227+
// The exclusion is not a detail. This project asks for comments that say WHY,
228+
// and a size limit counting them would be a limit on explaining - so the two
229+
// rules would pull against each other, and the one with a guard would win.
230+
// That is the shape docs/QUALITY.md names when it says the ceiling is measured
231+
// without comments.
232+
//
233+
// The measure itself had no test until 2026-08-05, so nothing would have
234+
// noticed the exclusion breaking. It would not have surfaced as a wrong number
235+
// either - it would have surfaced as a function suddenly over the ceiling for
236+
// having been explained, which reads like a real finding.
237+
//
238+
// Written the way the rule asks for: add N lines of comment and N lines of
239+
// code to the same function, and require that only the second moves the number.
240+
func TestTheSizeCeilingCountsCodeAndNotExplanation(t *testing.T) {
241+
measure := func(t *testing.T, source string) int {
242+
t.Helper()
243+
fset := token.NewFileSet()
244+
file, err := parser.ParseFile(fset, "measured.go", source, parser.ParseComments)
245+
if err != nil {
246+
t.Fatalf("parsing the sample: %v", err)
247+
}
248+
return codeLines(strings.Split(source, "\n"), commentLines(fset, file), 1, len(strings.Split(source, "\n")))
249+
}
250+
251+
const bare = `package sample
252+
253+
func f() {
254+
a := 1
255+
b := 2
256+
_ = a + b
257+
}
258+
`
259+
// Six lines of comment and a blank line, none of which is code.
260+
const explained = `package sample
261+
262+
// One.
263+
// Two.
264+
// Three.
265+
/* Four.
266+
Five.
267+
Six. */
268+
269+
func f() {
270+
a := 1
271+
b := 2
272+
_ = a + b
273+
}
274+
`
275+
// The same as bare, plus three lines that really are code.
276+
const bigger = `package sample
277+
278+
func f() {
279+
a := 1
280+
b := 2
281+
c := 3
282+
d := 4
283+
e := 5
284+
_ = a + b + c + d + e
285+
}
286+
`
287+
288+
base := measure(t, bare)
289+
if base == 0 {
290+
t.Fatal("the sample measured zero lines of code, so this guard would pass on anything")
291+
}
292+
293+
if n := measure(t, explained); n != base {
294+
t.Errorf("adding six lines of comment moved the measure from %d to %d.\n"+
295+
"The ceiling would then be a limit on explaining, which is the other rule this project runs on.",
296+
base, n)
297+
}
298+
if n := measure(t, bigger); n != base+3 {
299+
t.Errorf("adding three lines of code moved the measure from %d to %d, and %d was expected.\n"+
300+
"A measure that does not move for code is not measuring size at all.",
301+
base, n, base+3)
302+
}
303+
t.Logf("%d lines of code, unchanged by six lines of comment, plus three for three lines of code", base)
304+
}

‎internal/guard/crowding_test.go‎

Lines changed: 120 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,120 @@
1+
package guard
2+
3+
import (
4+
"fmt"
5+
"go/ast"
6+
"go/parser"
7+
"go/token"
8+
"os"
9+
"path/filepath"
10+
"sort"
11+
"strings"
12+
"testing"
13+
)
14+
15+
// The second knob of the ratchet: how many things are CROWDING the ceiling.
16+
//
17+
// The ceiling in codeshape_test.go answers one question - is anything over the
18+
// line - and it is blind to the shape that actually happens: nothing over the
19+
// line, and everything creeping towards it. A tree where thirty functions sit
20+
// at seventy nine lines passes that guard and is exactly the tree the guard
21+
// exists to prevent.
22+
//
23+
// So the ratchet needs two knobs rather than one, which is the general form of
24+
// "the threshold only goes down" applied to any metric that decays slowly:
25+
//
26+
// a ceiling on the worst single case - catches one thing growing to a record
27+
// a count of things near the ceiling - catches everything drifting at once
28+
//
29+
// Measured on 2026-08-05 before the numbers below were chosen, because a
30+
// threshold picked out of the air is a guess, and a guess written into a gate
31+
// is a guess nobody can argue with later.
32+
const (
33+
// What counts as crowding. Three quarters of the ceiling is far enough from
34+
// it that ordinary code does not trip the count, and close enough that
35+
// something arriving there is on its way.
36+
crowdingShare = 75
37+
38+
// Caps measured on the tree of 2026-08-05, then frozen - eleven functions and
39+
// two files were already crowding, against a first guess of four. That gap is
40+
// the whole argument for measuring rather than choosing: a cap of four would
41+
// have gone in red and been raised to make it pass, which is how a ratchet
42+
// becomes a rubber band. Like the ceilings
43+
// themselves these only go down. Raising one to turn a run green is the
44+
// same act as editing a golden value for the same reason.
45+
crowdedFunctions = 11
46+
crowdedFiles = 2
47+
)
48+
49+
func TestNothingIsQuietlyCreepingTowardsTheCeiling(t *testing.T) {
50+
var functions, files []string
51+
52+
for _, p := range packages(t) {
53+
for _, path := range p.files {
54+
body, err := os.ReadFile(path)
55+
if err != nil {
56+
t.Fatalf("reading %s: %v", path, err)
57+
}
58+
src := strings.Split(strings.ReplaceAll(string(body), "\r\n", "\n"), "\n")
59+
60+
fset := token.NewFileSet()
61+
file, err := parser.ParseFile(fset, path, body, parser.ParseComments)
62+
if err != nil {
63+
t.Fatalf("parsing %s: %v", path, err)
64+
}
65+
comments := commentLines(fset, file)
66+
67+
rel, err := filepath.Rel(repoRoot(t), path)
68+
if err != nil {
69+
rel = path
70+
}
71+
rel = filepath.ToSlash(rel)
72+
73+
if n := codeLines(src, comments, 1, len(src)); crowding(n, longestFile) {
74+
files = append(files, fmt.Sprintf("%s %d/%d", rel, n, longestFile))
75+
}
76+
77+
for _, decl := range file.Decls {
78+
fn, ok := decl.(*ast.FuncDecl)
79+
if !ok || fn.Body == nil {
80+
continue
81+
}
82+
from := fset.Position(fn.Pos()).Line
83+
to := fset.Position(fn.End()).Line
84+
if n := codeLines(src, comments, from, to); crowding(n, longestFunction) {
85+
functions = append(functions, fmt.Sprintf("%s:%d %s %d/%d",
86+
rel, from, name(fn), n, longestFunction))
87+
}
88+
}
89+
}
90+
}
91+
92+
sort.Strings(functions)
93+
sort.Strings(files)
94+
95+
if len(functions) > crowdedFunctions {
96+
t.Errorf("%d function(s) are within %d%% of the ceiling and the cap is %d:\n %s\n\n"+
97+
"Nothing is over the line, which is the point - this is the drift the other guard "+
98+
"cannot see. Split one of these rather than raising the cap.",
99+
len(functions), crowdingShare, crowdedFunctions, strings.Join(functions, "\n "))
100+
}
101+
if len(files) > crowdedFiles {
102+
t.Errorf("%d file(s) are within %d%% of the ceiling and the cap is %d:\n %s",
103+
len(files), crowdingShare, crowdedFiles, strings.Join(files, "\n "))
104+
}
105+
106+
t.Logf("crowding the ceiling: %d function(s) of %d allowed, %d file(s) of %d allowed",
107+
len(functions), crowdedFunctions, len(files), crowdedFiles)
108+
for _, f := range functions {
109+
t.Logf(" function %s", f)
110+
}
111+
for _, f := range files {
112+
t.Logf(" file %s", f)
113+
}
114+
}
115+
116+
// crowding reports whether a measurement has reached the share of the ceiling
117+
// at which it counts as on its way there.
118+
func crowding(n, ceiling int) bool {
119+
return n*100 >= ceiling*crowdingShare
120+
}

‎internal/guard/mutationcoverage_test.go‎

Lines changed: 2 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -87,6 +87,8 @@ var notProvenByMutation = map[string]bool{
8787
// "proven another way" are different states and lumping them together would
8888
// send a later session to re-prove what is already proven.
8989
var provenByProbe = map[string]string{
90+
"TestNothingIsQuietlyCreepingTowardsTheCeiling": "checked 2026-08-05 by hand. A function of seventy lines of code was dropped into internal/core, the count went from eleven to twelve against a cap of eleven, and the guard went red naming it. Removing the file put it back to green. It cannot be a mutation entry because the break is sixty extra lines rather than one substitution, and a mutation aimed at a function that happens to sit just under the threshold today would go stale at the first refactor that moves it.",
91+
9092
"TestEveryTextFormatIsValidUTF8": "checked 2026-08-02 with tools/probes/probe-utf8-filler.py, which swaps in a vocabulary of Polish words and sweeps 304 sizes. " +
9193
"Before core.AppendFiller cut on a character boundary: 304 files, every one the right size, 86 of them carrying invalid UTF-8. After: 0. " +
9294
"It was a mutation until the fix landed, and the runner then reported it NOT CAUGHT - which is the fix being proven rather than a hole. " +

‎internal/guard/presetwindow_test.go‎

Lines changed: 22 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -130,13 +130,30 @@ func TestThePresetScreenAndTheCommandLineProduceTheSameRun(t *testing.T) {
130130
cliNames, windowNames)
131131
}
132132

133+
// Two empty sets are equal, and an equality that can be satisfied by
134+
// emptiness is an equality that proves nothing. The comparison below walks
135+
// this list, so a run that produced nothing but a manifest would compare
136+
// zero files and pass - which is the shape that took another project's
137+
// guard from catching 30 injected faults out of 30 down to 13, with the
138+
// mutation report still looking clean.
139+
//
140+
// size-boundaries is seven files plus the manifest, and asserting the count
141+
// here rather than logging it is the whole difference.
142+
const wanted = 7
143+
if len(cliNames) != wanted+1 {
144+
t.Fatalf("the preset produced %d thing(s) and %d files plus a manifest was expected: %v",
145+
len(cliNames), wanted, cliNames)
146+
}
147+
133148
// Byte for byte, because "the same set" and "the same files" are different
134149
// claims and only the second one is worth anything to somebody whose test
135150
// suite hashes them.
151+
compared := 0
136152
for _, name := range cliNames {
137153
if name == "manifest.json" {
138154
continue
139155
}
156+
compared++
140157
a, err := os.ReadFile(filepath.Join(fromCLI, name))
141158
if err != nil {
142159
t.Fatalf("reading %s from the command line run: %v", name, err)
@@ -172,7 +189,11 @@ func TestThePresetScreenAndTheCommandLineProduceTheSameRun(t *testing.T) {
172189
t.Errorf("the recipe hashes differ, so the two surfaces expanded the preset differently.\n"+
173190
" command line: %s\n window: %s", a, b)
174191
}
175-
t.Logf("%d file(s) identical, one preset block, one recipe hash", len(cliNames)-1)
192+
if compared != wanted {
193+
t.Errorf("%d file(s) were compared and %d were expected - an equality nothing was measured against",
194+
compared, wanted)
195+
}
196+
t.Logf("%d file(s) identical, one preset block, one recipe hash", compared)
176197
}
177198

178199
// A parameter the caller left alone is recorded as ours rather than theirs.

‎internal/guard/textformats_test.go‎

Lines changed: 51 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -431,11 +431,59 @@ func TestAnSVGDrawingCarriesRealShapes(t *testing.T) {
431431
}
432432
}
433433

434-
// textFormats are the formats whose whole output is text. A new one belongs
435-
// here, and the check below fails on a name that is not a format at all, so the
436-
// list cannot quietly point at nothing.
434+
// textFormats are the formats whose whole output is text, and binaryFormats is
435+
// everything else. Between them they have to name every registered format.
436+
//
437+
// Two lists rather than one, and that is the repair. The check below fails on a
438+
// name that is not a format, so neither list can point at nothing - but until
439+
// 2026-08-05 nothing asked the other direction, and that direction is the one
440+
// that matters. A format registered tomorrow simply would not have been checked
441+
// for valid UTF-8, and the suite would have stayed green while covering less.
442+
//
443+
// This is the shape docs/OBSERVATIONS.md now calls out: an audit of
444+
// completeness has to run FROM THE SOURCE towards the list. Walking the entries
445+
// already written down cannot, by construction, find what is missing from them.
437446
var textFormats = []string{"txt", "md", "log", "csv", "json", "xml", "html", "svg"}
438447

448+
var binaryFormats = []string{"pdf", "png", "targz", "wav", "zip"}
449+
450+
// Every registered format is on exactly one of the two lists above.
451+
//
452+
// Cheap, and it is the only thing standing between "eight formats are checked
453+
// for valid UTF-8" and "eight of the formats that existed when somebody last
454+
// looked". Adding a format now forces the question rather than skipping it in
455+
// silence.
456+
func TestEveryFormatIsClassifiedAsTextOrBinary(t *testing.T) {
457+
said := map[string]string{}
458+
for _, id := range textFormats {
459+
said[id] = "text"
460+
}
461+
for _, id := range binaryFormats {
462+
if kind, twice := said[id]; twice {
463+
t.Errorf("%s is on both lists, so nobody can tell which it is (%s)", id, kind)
464+
}
465+
said[id] = "binary"
466+
}
467+
468+
registered := map[string]bool{}
469+
for _, id := range format.IDs() {
470+
registered[id] = true
471+
if said[id] == "" {
472+
t.Errorf("the registry has %s and neither list names it, so nothing says whether its "+
473+
"output has to be valid UTF-8. Put it on textFormats or on binaryFormats.", id)
474+
}
475+
}
476+
for id := range said {
477+
if !registered[id] {
478+
t.Errorf("%s is classified and the registry no longer has it - remove the entry", id)
479+
}
480+
}
481+
if len(registered) == 0 {
482+
t.Fatal("no format was registered, so this guard would pass without checking anything")
483+
}
484+
t.Logf("%d format(s): %d text, %d binary", len(registered), len(textFormats), len(binaryFormats))
485+
}
486+
439487
// Every record format pads its last value to an exact BYTE count and then cuts
440488
// to length. Today every word in every vocabulary is ASCII, so a byte is a
441489
// character and the cut always lands between two of them.

0 commit comments

Comments
 (0)